About Turing
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L
Role Overview
We are seeking skilled Terminal-Bench Task 2.0 (Harbor) Authors to design, develop, and validate high-quality benchmark tasks for evaluating large language models (LLMs) in simulated Linux terminal environments.
In this role, you will create challenging, deterministic, and reproducible tasks that rigorously test AI capabilities across software engineering, systems, data science, and mathematical domains. Your work will directly contribute to evaluating and stress-testing frontier AI models under real-world terminal constraints.
What does a typical day look like?
- Author original Terminal-Bench 2.0 (Harbor) tasks with precise, unambiguous instructions
- Design realistic Linux terminal workflows involving filesystems, processes, networking, and containers
- Implement golden solutions and pytest-based evaluation scripts with deterministic outcomes
- Build and maintain Dockerized environments with pinned dependencies for reproducibility
- Anticipate edge cases and failure modes to prevent reward hacking
- Validate tasks by running them against frontier LLMs and iterating to achieve target pass/fail rates
- Deliver approximately 5 fully validated benchmark tasks per week
Key Responsibilities
- Design & Author Tasks: Create unique, high-quality Terminal-Bench 2.0 tasks with clear goals, environment setup, and expected outputs
- Write Deterministic Tests: Develop robust pytest-based test suites with no hidden test cases
- Build Docker Environments: Configure reproducible containerized setups with locked dependencies
- Model Validation: Execute tasks against AI models and refine difficulty and clarity
- Quality Assurance: Ensure strong alignment between instructions, tests, and evaluation logic to maintain benchmark integrity
Required Skills & Qualifications
Technical Skils
- Python: Strong proficiency with clean, testable code (3+ years experience)
- Bash / Shell Scripting: Confident with Unix command-line tools and workflows (2+ years experience)
- Linux Systems: Familiarity with filesystems, permissions, processes, and basic networking (1+ year experience)
- Docker: Experience building, configuring, and debugging containerized environments
Domain Expertise (one or more required)
- Software Engineering, System Administration, or Debugging
- Data Science, Machine Learning, or Model Training
- Mathematics, Algorithm Design, or Scientific Computing
- Frontend or Backend Development
- Data Preprocessing and Analysis
Core Competencies
- Strong analytical and problem-solving ability
- Extreme attention to detail in technical writing
- Ability to anticipate corner cases and unintended model behaviors
- Comfort working with deterministic evaluation and strict correctness criteria
Expected Output
- ~5 fully validated Terminal-Bench tasks per week, including instructions, Docker setup, golden solutions, and test suites
Perks of Freelancing With Turing:
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
Offer Details
- Commitments Required: 8 hours per day with overlap of 4 hours with PST.
- Employment type : Contractor assignment (no medical/paid leave)
- Duration of contract : 1 month; [expected start date is next week]
- Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Mexico