Go back to all jobs

Senior AI QA Engineer

Full-time
On-site
Skills
QA
Software Quality Assurance
Overview

Job Description – Senior AI Quality Assurance (QA) Engineer

Domain: Banking, Financial Services & Insurance (BFSI)

Experience: 7+ Years
Location: Bengaluru
Work Mode: Hybrid
Employment Type: Full-Time

About the Role

We are looking for an experienced Senior AI Quality Assurance (QA) Engineer to ensure the quality, reliability, and production readiness of enterprise-grade AI applications.

This role extends beyond traditional software testing and focuses on the evaluation and quality engineering of LLM applications, RAG systems, and AI agents. You will own automated testing, AI evaluations, benchmark creation, observability, prompt regression testing, and end-to-end validation of AI workflows. Working closely with AI Engineers, Product Managers, and Platform teams, you will establish measurable quality standards and ensure every release meets enterprise-grade expectations for accuracy, reliability, performance, and scalability.

Key Responsibilities

AI Evaluation & Benchmarking

Design benchmark (golden) datasets and build automated evaluation pipelines for LLM applications. Define quality gates and continuously evaluate prompts, models, retrieval pipelines, and agent behaviour using metrics such as hallucination rate, tool selection accuracy, execution accuracy, precision, recall, latency, and cost.

RAG & Agent Quality Validation

Validate Retrieval-Augmented Generation (RAG) pipelines and AI agent workflows by testing retrieval quality, context relevance, tool invocation, reasoning flow, memory, and end-to-end task completion. Design evaluation scenarios covering ambiguous queries, multi-turn conversations, retrieval failures, and edge cases.

Python Automation & API Testing

Develop and maintain scalable automation frameworks using Python and Pytest for unit testing, integration testing, API testing, regression testing, and end-to-end validation. Build reusable test utilities and integrate automated quality checks into CI/CD pipelines.

Frontend Automation

Develop automated UI test suites using Playwright to validate AI-powered user journeys, conversational interfaces, workflow execution, and end-to-end application behaviour across releases.

Observability & Root Cause Analysis

Use OpenTelemetry, tracing platforms, and AI observability tools to analyse execution traces, latency, model responses, API calls, and workflow behaviour. Perform root cause analysis to identify regressions, hallucinations, bottlenecks, and production issues.

Performance & Enterprise Readiness

Validate AI application performance by monitoring latency, throughput, reliability, and scalability. Ensure production readiness through regression testing, API validation, workflow testing, and enterprise quality standards.

Required Skills

  • 7+ years of experience in Software QA, Test Automation, or AI Quality Engineering.
  • Strong Python programming skills.
  • Hands-on experience with Pytest for:
    • Unit testing
    • Integration testing
    • API testing
    • Regression testing
  • Experience testing REST APIs and backend services.
  • Experience evaluating LLM-powered applications.
  • Hands-on experience evaluating Retrieval-Augmented Generation (RAG) systems.
  • Understanding of AI agent evaluation methodologies.
  • Experience measuring AI quality using metrics such as:
    • Hallucination Rate
    • Tool Selection Accuracy
    • Execution Accuracy
    • Precision / Recall
    • Latency (P50/P95/P99)
    • Token Usage
  • Experience with Playwright or similar frontend automation frameworks.
  • Experience with OpenTelemetry, tracing, or observability platforms.
  • Strong debugging and root cause analysis skills.
  • Experience integrating automated tests into CI/CD pipelines.

Nice to Have

  • Experience with LangSmith, Langfuse, MLflow, Arize Phoenix, or similar AI observability platforms.
  • Experience evaluating multi-agent systems and orchestration frameworks such as LangGraph, CrewAI, Google ADK or AutoGen.
  • Experience with vector databases such as Pinecone, Milvus, Weaviate, pgvector, or Vertex AI Vector Search.
  • Exposure to OpenAI, Anthropic, Gemini, or Azure OpenAI.
  • Experience with performance and load testing tools.
  • Prior experience in Banking, Financial Services, or Insurance (BFSI).

Education

BE / BTech / MCA / MTech / BSc in Computer Science, Artificial Intelligence, Data Science, Information Technology, or a related field.

Success in this Role

Success in this role will be measured by your ability to build a scalable AI quality engineering framework through automated testing, benchmark-driven evaluations, RAG and agent validation, frontend and API automation, and comprehensive observability. You will help ensure our AI applications consistently deliver accurate, reliable, performant, and production-ready experiences for enterprise users.

Turing
Create an account
Already have an account?
Or continue with email
Trusted by AI leaders, enterprises, and more
Anthropic
Dell
Disney
Nvidia
Pepsi
Reddit
Rivian
snowflake
Anthropic
Dell
Disney
Nvidia
Pepsi
Reddit
Rivian
snowflake
Terms of ServicePrivacy Policy© 2026 Turing Enterprises, Inc.

Don't miss out on this job opportunity!

Senior AI QA Engineer