About Turing:
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview:
We are staffing a frontier AI data initiative that requires strong software engineers to build the infrastructure and training data used to develop and evaluate AI agents.
The work spans two closely related tracks: Connectors and Tasks. Candidates may specialize in one track, but the strongest profiles are generalist backend engineers who can contribute across both.
What you'll do:
Build Python backend applications that faithfully replicate real-world SaaS tools such as Slack, Linear, Jira, Notion, Gmail, wikis, and other enterprise applications.
Develop realistic digital-work environments where AI agents can interact with tools, data, and workflows.
Mine real-world data, tools, and workflows to understand how knowledge work is performed across enterprise applications.
Author realistic, long-horizon agent tasks that require multi-step reasoning, tool usage, and interaction with connector environments.
Write clear evaluation rubrics that precisely define correct, complete, and high-quality task completion.
Validate tasks end-to-end for quality, realism, correctness, and achievability through rigorous QA.
Identify ambiguity, edge cases, grading gaps, and unrealistic assumptions before tasks ship.
Reproduce and debug failures across tasks and connector environments and provide actionable feedback to other engineers.
Use AI coding tools extensively throughout development, debugging, testing, and QA workflows.
What we're looking for:
Strong backend software engineering experience, primarily in Python.
Experience building, testing, debugging, and maintaining backend applications.
Working knowledge of GCP, Docker, virtual machines, and Harbor.
Strong engineering judgment and a high bar for correctness, with the ability to identify edge cases and ambiguous requirements.
Ability to understand real-world workflows and translate them into precise technical implementations, tasks, and evaluation criteria.
Excellent written communication skills, particularly for technical documentation, QA feedback, and rubric writing.
High daily proficiency with AI coding tools such as Claude Code, Cursor, Copilot, or similar tools — this is a hard requirement.
Ability to work independently across unfamiliar codebases, tools, and technical environments.
Nice to have:
Experience building or evaluating AI agents, LLM applications, or agentic systems.
Experience developing backend integrations or applications that replicate or interact with SaaS products.
Familiarity with enterprise SaaS tools such as Slack, Linear, Jira, Notion, Gmail, or wikis at an API or data-model level.
Experience with evaluation design, QA, rubric development, or AI training-data generation.
Experience working with complex multi-step workflows involving tool/function calling.
Ability to contribute across both Connectors and Tasks, rather than specializing exclusively in one track.
Offer Details:
Commitments Required: At least 6 hours per day and minimum 40 hours per week with overlap of 6 hours with PST.
Employment type : Contractor assignment (no medical/paid leave)
Duration of contract : 5 week [expected start date is next week]
Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Mexico
Don't miss out on this job opportunity!
Software Engineer Pod Lead