About Turing
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview
We are seeking an AI Training & Evaluation Specialist (Biology/Health) to design, curate, and review advanced biological and health science assessment tasks to train and evaluate state-of-the-art AI models. This dual-focus role involves two key work streams: authoring complex, text-only biology problems from graduate-level concepts and reviewing domain-specific question-answering tasks for scientific rigor and quality. You will play a vital role in identifying failure modes in frontier AI systems, ensuring generated outputs meet high academic and clinical standards.
Requirements
- Education: Master’s degree or PhD in Biology, Health Sciences, Medicine, or a closely related biological field.
- Domain Expertise: Strong mastery of upper-undergraduate and graduate-level biological systems, health concepts, and clinical or research-based reasoning.
- Writing & Scientific Translation: Fluent written English with the ability to write precise, rigorous scientific explanations and convert visual/diagram-heavy concepts into self-contained, text-only problems.
- Evaluation Skills: Ability to critically navigate domain topics, judge the quality and accuracy of complex questions/answers, diagnose AI failure modes, and refine tasks to meet high difficulty targets.
- Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable).
- Technical Infrastructure: Personal desktop/laptop equipped with a stable, high-speed internet connection in a remote setup.
Responsibilities
- Problem Design & Curation: Author and curate upper-undergraduate and graduate-level Biology problems with clear, fully worked solutions, converting diagram-heavy or multi-part questions into text-only, solvable tasks.
- Quality Review & Audit: Review, evaluate, and judge biology and health question-answering tasks created by others or generated by AI models, ensuring scientific correctness, terminology precision, and clarity.
- AI Model Evaluation: Diagnose AI-generated errors, expose model failure modes, and refine questions to calibrate target difficulty levels.
- Collaborative Quality Control: Collaborate with cross-functional teams to maintain consistent task standards across large datasets and peer-review tasks prior to final delivery.
Education & Experience
- Prior experience in teaching, grading, university-level exam creation, or research-based analytical work required.
- Familiarity with AI training data best practices is preferred; 1+ month of prior experience on Turing AI evaluation projects is a strong plus (not mandatory).
Offer Details:
- Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.
- Engagement type: Contractor
- Engagement Length: 4 weeks
Evaluation Process -
- Shortlisted candidates will be sent a Job Interest Form.
- Final selected candidates will be contacted with the next steps.