About Turing:
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview:
We're staffing a frontier AI data initiative that builds the training data and evaluations used to develop and measure AI agents. You'll join the Mining team, whose job is to turn real-world work into rigorous, long-horizon tasks that AI agents can be trained and tested against.
What you'll do:
- Mine real data, tools, and workflows to understand how knowledge work actually gets done across enterprise applications.
- Author long-horizon agent tasks grounded in that research — realistic, multi-step objectives that mirror genuine digital work.
- Write clear evaluation rubrics that define what correct, complete, and high-quality completion looks like.
- Validate task quality, realism, and correctness through rigorous QA, catching ambiguity, unrealistic assumptions, and grading gaps before tasks ship.
- Work closely with the connectors team — engineers who build Python backends that faithfully replicate SaaS tools (Slack, Linear, Jira, Notion, Gmail, wikis, and similar) — to ground your tasks in realistic environments.
What we're looking for:
- Strong backend software engineering, primarily in Python.
- Working knowledge of GCP, Docker, virtual machines, and Harbor.
- Sound engineering judgment and a high bar for correctness — you find edge cases and ambiguity others miss.
- Ability to reason about real-world workflows and translate them into precise, well-scoped tasks and rubrics.
- Excellent written communication; rubric and QA writing is core to the role.
- High daily proficiency with AI coding tools (e.g., Claude Code, Cursor, Copilot) — a hard requirement, not a nice-to-have.
Nice to have:
- Experience building or evaluating agentic/LLM systems.
- Familiarity with the SaaS tools above at an API/data-model level.
- Background in data annotation, evaluation design, or QA at scale.
- Candidates may specialize in mining, but the strongest profiles are excellent generalist backend engineers who can also contribute on the connectors side when needed.
- Review long-horizon agent tasks authored by the mining team for realism, correctness, and scope — catching flawed assumptions, ambiguity, and missing edge cases before tasks ship.
- Pressure-test rubrics: confirm they define completion precisely, grade consistently, and can't be gamed or misread.
- Validate tasks end-to-end against the connector environments they run in, verifying that stated objectives are actually achievable and that expected outcomes hold.
- Reproduce and debug failures, then feed clear, specific findings back to task authors and connector engineers.
- Define and improve QC standards, checklists, and processes as the initiative scales.
Offer Details:
- Commitments Required: At least 6 hours per day and minimum 40 hours per week with overlap of 6 hours with PST.
- Employment type : Contractor assignment (no medical/paid leave)
- Duration of contract : 5 week [expected start date is next week]
- Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Mexico