About Turing:
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.
Role Overview:
We are seeking US-based generalist raters to participate in Gemini Personalization Evals. In this role, raters will perform generalist evaluation work using personalized prompts designed to retrieve context and data across connected Google apps.
Key Qualifications
- Location: Must be currently based in the United States.
- App Connectivity: Must be willing and able to connect their Google apps (e.g., Gmail, Google Photos, Google Calendar, Google Drive, etc.) to Gemini to support data retrieval and personalized prompt evaluations.
- Role Scope: Perform generalist rating and evaluation tasks, assessing model performance, context retrieval, and response quality across connected apps.
- Account Activity: Active use of standard Google workspace/consumer apps with sufficient personal data/history to generate and evaluate retrieval prompts.
- Exceptional Analytical Thinking: Ability to evaluate nuanced AI responses and identify strengths and weaknesses in personalization quality.
- Attention to Detail: Strong ability to review AI-generated responses and identify subtle issues or inconsistencies.
- Communication: Excellent written communication skills for providing clear and structured evaluation feedback.
- Independence: Self-motivated and able to work independently in a remote environment.
- Technical Setup: Desktop or laptop with a reliable internet connection.
Description:
In this role, you will be part of a team responsible for evaluating the quality of personalized AI interactions for business users. Your day-to-day responsibilities will include:
- Evaluating AI-generated responses based on your professional email history and business account activity.
- Assessing whether responses are relevant, accurate, helpful, and appropriately personalized.
- Identifying issues related to incorrect personalization, unsupported assumptions, or irrelevant recommendations.
- Comparing multiple AI responses and determining which provides the better overall user experience.
- Providing clear, detailed, and well-reasoned evaluation feedback to support model improvements.
- Following project guidelines to ensure consistent, accurate, and high-quality evaluations.
- Maintaining data privacy and adhering to all project confidentiality requirements throughout the evaluation process.
Education & Experience
- Bachelor's degree or equivalent practical experience in any field.
- Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required.
Offer Details:
- Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.
- Engagement type: Contractor
- Engagement Length: upto 4 months
Evaluation Process -
- Shortlisted candidates will be sent a Job Interest Form.
- Final selected candidates will be contacted with the next steps, including consent and onboarding requirements.
Consent & Compliance:
Must be willing to grant the necessary app integration permissions required for the project and complete all standard consent forms and non-disclosure agreements.