AI Evaluation Engineer — Evals & Systems Verification

Company: Fermi AI

Location: Bangalore, India

Type: Full-time

Posted: 2026-08-03

About this role

AI Evaluation Engineer — Evals & Systems Verification

Location: Bengaluru (Hybrid)

Experience: 3–6 years (indicative)

Start: In 2–3 months

### About the Role

Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.

### Responsibilities

  • Build and maintain evaluation harnesses for AI-facing features to measure and tune system quality (e.g., capture quality, retrieval quality, guidance quality)
  • Own end-to-end and API test infrastructure (Playwright-class), supporting a continuous, daily-release cycle
  • Design and execute host-behavior probes using scripted user sessions across diverse AI assistants, ensuring product behavior aligns with expectations
  • Gate production releases through thorough user acceptance testing (UAT), regression analysis, and quality reporting

### Requirements

  • Background in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure—not just executing tests
  • Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress
  • Familiarity with LLM applications or a demonstrated interest in evaluating non-deterministic systems
  • Highly autonomous and able to define your own workflows and processes for evaluation and verification

Create Your Job Alert

Other AI Jobs

Other Jobs in Bangalore