Go back to all jobs

LLM Python Reviewer

Full-time
Remote
Skills
Python
Overview

About Turing

Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L


Role Overview

Own end-to-end quality and calibration for LLM/agent evaluations at Turing. You’ll lead prompt/rubric governance, analyze agent trajectories, root-cause failures, and ensure consistent signal across evaluators and datasets—using strong Python+SQL skills and advanced prompt engineering.


What does day-to-day life look like?

  • Define, run, and improve evaluation pipelines for LLMs/agents (prompt/rubric/verifier governance).
  • Use Python for evaluation, analysis, and review workflows; use SQL to query/audit results and detect drift.
  • Investigate agent trajectories to identify failure modes (hallucinations, tool misuse, systematic errors) and drive fixes.
  • Calibrate evaluators/datasets; establish consistency standards and reviewer guides.
  • Partner with product/research to design experiments, monitor metrics, and ship improvements.
  • Contribute to agentic tool-calling and MCP environment evaluations and best practices.
  • Communicate findings and recommendations clearly to stakeholders.

Requirements

  • 3–5+ years relevant experience; ≥1.5 years at Turing.
  • Prior Pod Lead or Calibrator experience at Turing.
  • Strong Python proficiency (evaluation, analysis, review automation).
  • SQL experience for querying and auditing evaluation outputs.
  • Advanced prompt engineering for LLM/agent systems; experience reviewing/approving prompts, verifiers, rubrics.
  • Deep understanding of LLM/agent behavior and failure modes (hallucination, tool misuse, systematic errors).
  • Proven ability to analyze agent trajectories and determine root causes.
  • Track record ensuring calibration/consistency across evaluators and datasets.
  • Hands-on with agentic tool calling and MCP environments.
  • Strong analytical judgment; excellent written and verbal English communication.

Perks of Freelancing With Turing

  • Work in a fully remote environment.
  • Opportunity to work on cutting-edge AI projects with leading LLM companies.

Offer Details

  • Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. 
  • Engagement Type: Contractor assignment (no medical/paid leave)
  • Duration of Contract: 3 months (adjustable based on engagement)
  • Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, Mexico

Evaluation Process

  • Two rounds of technical interviews
Turing
Create an account
Already have an account?
Or continue with email
Trusted by AI leaders, enterprises, and more
Anthropic
Dell
Disney
Nvidia
Pepsi
Reddit
Rivian
snowflake
Anthropic
Dell
Disney
Nvidia
Pepsi
Reddit
Rivian
snowflake
Terms of ServicePrivacy Policy© 2026 Turing Enterprises, Inc.

Don't miss out on this job opportunity!

LLM Python Reviewer