About Turing
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L
Role Overview
We’re seeking a hands-on Quality Engineering leader to define, own, and execute the end-to-end quality strategy for full-stack and LLM-powered products. You’ll lead quality initiatives, build automation and evaluation pipelines in Python/Node.js, and ensure release readiness through clear gates, data-driven evaluations, and robust CI/CD integration.
What does day-to-day life look like?
- Own quality strategy for complex, distributed systems; set quality bars, processes, and metrics (SLOs, DORA/defect escape, test coverage).
- Design and implement automation frameworks and evaluation pipelines in Python and Node.js.
- Drive testing across layers: E2E, Frontend/UI, Backend/API, and Performance (load & stress).
- Build and operate quality systems for LLM products:
AI agents and multi-step agent workflows.
MCP tools/servers/integrations.
LLM-as-a-Judge and automated evaluation harnesses. - Define golden datasets, regression prompts, judge calibration, and handle non-deterministic/probabilistic behavior validation.
- Integrate quality into CI/CD with release gates, test orchestration, and dashboards; own release readiness.
- Lead root-cause analysis, quality reviews, and continuous improvement initiatives.
- Mentor QA engineers/SDETs, clarify ownership boundaries, and partner with Engineering, Product, and Research.
Requirements
- 7+ years in Quality Engineering/SDET/Software Engineering roles.
- 2–3+ years leading quality initiatives or owning quality strategy for complex systems, including hands-on QA evaluation.
- Strong development skills in Python and Node.js (test frameworks, eval pipelines, automation tooling).
- Proven experience defining and owning end-to-end quality for full-stack applications.
- Ownership of testing across E2E, UI, Backend/API, Performance (load & stress).
- 2+ years building automated quality/evaluation systems for LLM-powered products, including agents/workflows, MCP components, and LLM-as-a-Judge pipelines.
- Experience with golden datasets, regression prompts, and judge calibration; adept at validating non-deterministic systems.
- CI/CD integration, release gates, and release readiness ownership.
- Experience mentoring QA/SDET teams and defining clear responsibility boundaries.
- Excellent written and verbal communication in English.
Perks of Freelancing With Turing
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
Offer Details
- Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.
- Engagement Type: Contractor assignment (no medical/paid leave)
- Duration of Contract: 3 months (adjustable based on engagement)
- Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, Mexico
Evaluation Process
- Two rounds of technical interviews