Senior AI Agent Quality Engineer

MultiPlan

McLean (VA)

On-site

USD 140,000 - 160,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401k
Bonus opportunity

Job summary

Claritev in McLean, Virginia, is seeking a Senior AI Operations Engineer to serve as the quality gatekeeper for agentic AI systems powering our healthcare products. You will design and run test suites, benchmarks, and evaluation pipelines to validate agent behavior before and after release.

Collaborate with AI engineers, data scientists, and product teams to build golden datasets, automate testing in CI/CD, triage and reproduce failures, and translate evaluation results into actionable quality

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Data Science, or a related quantitative field.
  • Master's degree is a plus.
  • 5+ years hands-on experience in software engineering, QA/test automation, ML engineering, or related technical discipline.
  • 2+ years testing/operating production ML/AI systems.
  • 1+ years hands-on with generative AI, LLMs, RAG, or agentic AI systems.
  • Track record building test automation or evaluation capabilities.

Responsibilities

  • Design, build, and maintain automated test suites and evaluation pipelines for agentic AI systems, covering single-turn, multi-turn, and end-to-end workflow scenarios.
  • Develop golden datasets, test scenarios, and simulation environments reflecting real healthcare workflows and edge cases.
  • Execute and automate agent benchmarking, including regression testing across model, prompt, and tool changes, with clear pass/fail quality gates in CI/CD.
  • Measure and report agent quality metrics: task completion, tool-call accuracy, grounding/hallucination rates, response quality, latency, cost per task, and safety compliance.
  • Implement LLM-as-judge and rubric-based scoring approaches, and validate automated scores against human review.
  • Triage, reproduce, and root-cause agent failures by analyzing traces of LLM calls, tool invocations, retrieval steps, and orchestration paths; file actionable defect reports.
  • Perform adversarial and negative testing, including prompt-injection probes, malformed inputs, boundary conditions, and guardrail verification.
  • Monitor production agent behavior, investigate quality incidents and drift, and feed findings back into offline test coverage.
  • Contribute to evaluation dashboards, runbooks, and quality documentation used by engineering and product teams.
  • Verify PHI/PII handling, auditability, and human-in-the-loop controls behave as designed, supporting HIPAA and data-governance compliance.
  • Collaborate with AI engineers and data scientists to make agents more testable, observable, and reliable by design.

Skills

Python
pytest
APIs
distributed systems
LLM evaluation
LangChain
Kubernetes
CI/CD
observability tooling
statistics
guardrails

Education

Bachelor's degree in Computer Science, Engineering, Data Science, a quantitative discipline, or a related field
Master's degree a plus

Tools

LangSmith
Langfuse
Arize Phoenix
Ragas
DeepEval
promptfoo
LangGraph
LangChain
AutoGen
CrewAI

Job description

We are seeking a Senior AI Operations Engineer to serve as the quality gatekeeper for the agentic AI systems powering Claritev's next generation of healthcare products.

This is a hands-on Agentic QA role for an engineer who is passionate about answering a deceptively hard question: how do you prove an AI agent works? You will design and run the test suites, benchmarks, and evaluation pipelines that validate agent behavior before and after release - catching regressions, hallucinations, broken tool calls, and unsafe actions before they reach production healthcare workflows.

You will work closely with our Principal Agentic AI Operations Engineer, AI engineering teams, and Product stakeholders to build golden datasets, automate agent testing in CI/CD, triage and reproduce agent failures, and turn evaluation results into actionable quality insights. This role is an excellent path for a strong QA/test automation, MLOps, or backend engineer looking to specialize in the fast-growing discipline of AI agent quality and evaluation.

Job Roles and Responsibilities
  • Design, build, and maintain automated test suites and evaluation pipelines for agentic AI systems, covering single-turn, multi-turn, and end-to-end workflow scenarios.
  • Develop and curate golden datasets, test scenarios, and simulation environments that reflect real healthcare workflows and edge cases.
  • Execute and automate agent benchmarking, including regression testing across model, prompt, and tool changes, with clear pass/fail quality gates in CI/CD.
  • Measure and report agent quality metrics: task completion, tool-call accuracy, grounding/hallucination rates, response quality, latency, cost per task, and safety compliance.
  • Implement LLM-as-judge and rubric-based scoring approaches, and validate automated scores against human review.
  • Triage, reproduce, and root-cause agent failures by analyzing traces of LLM calls, tool invocations, retrieval steps, and orchestration paths; file actionable defect reports.
  • Perform adversarial and negative testing, including prompt-injection probes, malformed inputs, boundary conditions, and guardrail verification.
  • Monitor production agent behavior, investigate quality incidents and drift, and feed findings back into offline test coverage.
  • Contribute to evaluation dashboards, runbooks, and quality documentation used by engineering and product teams.
  • Verify PHI/PII handling, auditability, and human-in-the-loop controls behave as designed, supporting HIPAA and data-governance compliance.
  • Collaborate with AI engineers and data scientists to make agents more testable, observable, and reliable by design.
Job Requirements
Education
  • Bachelor's degree in Computer Science, Engineering, Data Science, a quantitative discipline, or a related field required.
  • Master's degree a plus.
Experience
  • 5+ years of hands-on experience in software engineering, QA/test automation engineering, ML engineering, or a related technical discipline.
  • 2+ years of experience testing, operating, or supporting production ML or AI systems.
  • 1+ years of hands-on experience with generative AI, LLMs, RAG, and/or agentic AI systems (professional projects preferred; substantial personal or open-source work considered).
  • Track record of building test automation or evaluation capabilities that measurably improved software or model quality.
Technical Skills
  • Strong Python skills, including test frameworks (e.g., pytest) and building automation/tooling; comfort working with APIs and distributed systems.
  • Hands-on experience with LLM/agent evaluation concepts: eval datasets, LLM-as-judge, rubric-based scoring, and regression detection.
  • Familiarity with evaluation and observability tooling such as LangSmith, Langfuse, Arize Phoenix, Ragas, DeepEval, promptfoo, or equivalent.
  • Working knowledge of agentic AI frameworks and patterns (e.g., LangGraph, LangChain, AutoGen, CrewAI), including tool use, orchestration, and guardrails.
  • Understanding of RAG pipelines, embeddings, and vector search, and how retrieval quality affects agent behavior.
  • Experience with CI/CD pipelines and integrating automated tests and quality gates into them.
  • Familiarity with cloud environments, containers, and log/trace analysis for debugging distributed systems.
  • Solid grasp of statistics fundamentals for interpreting evaluation results (sampling, variance, significance).
Other Skills
  • Strong analytical and debugging skills, with high attention to detail and a healthy skepticism toward "it works on my machine."
  • Clear written communication, especially in defect reports, test plans, and quality summaries.
  • Ability to operate effectively in a fast-moving, cross-functional environment.
Preferred Qualifications
  • Experience with Oracle Cloud Infrastructure (OCI), including its generative AI, data science, database, and observability capabilities.
  • Experience in healthcare, health technology, insurance, claims, payment integrity, or other regulated industries.
  • Experience testing systems that process sensitive data, including PHI or PII, in HIPAA-regulated environments.
  • Exposure to public agent/LLM benchmarks (e.g., SWE-bench, GAIA, AgentBench, tau-bench) or red-teaming and safety evaluation of LLM systems.
  • Experience with Kubernetes, infrastructure-as-code (e.g., Terraform), or SRE practices.
  • Prior QA/SDET background applied to ML or data-intensive systems.
Compensation

The salary range for this position is $140K-160K/year. Specific offers take into account a candidate's education, experience and skills, as well as the candidate's work location and internal equity. This position is also eligible for health insurance, 401k and bonus opportunity.

#LI-MC2

Why Claritev?
Healthcare is complex. We help make it clearer.

At Claritev, you'll do work that matters. Together, we're helping make healthcare more transparent and affordable for all through the power of data, technology, and expertise. We offer meaningful opportunities to grow your career, collaborate with talented colleagues, and make an impact on the clients and communities we serve. If you're looking for purpose, growth, and a team that succeeds together, you'll find it here.

What Guides Us

At Claritev, innovation, agility, and a focus on results drive our success. We embrace bold thinking, work as one team, take ownership, and strive for excellence in everything we do - creating meaningful impact for our clients, communities, and each other.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Agentic AI Operations Engineer
Principal Agentic AI Operations Engineer

MultiPlan • McLean (VA)

On-site
USD 165,000 - 185,000
Health insurance
401k
Bonus opportunity
Principal Applied AI Engineer, Agentic AI
Principal Applied AI Engineer, Agentic AI

MultiPlan • McLean (VA)

On-site
USD 190,000 - 210,000
Health insurance
401(k) retirement plan
Performance bonus
Senior Applied AI Scientist
Senior Applied AI Scientist

Claritev • United States

On-site
USD 120,000 - 150,000
Medical, dental, and vision coverage
401(k) with company match
Paid time off
QE Architect
QE Architect

MultiPlan • McLean (VA)

On-site
USD 90,000 - 145,000
Health insurance
401k
Bonus opportunity
Principal Engineer/Architect, AI Innovation
Principal Engineer/Architect, AI Innovation

Claritev • United States

On-site
USD 180,000 - 230,000
Health insurance
Bonus opportunity
401(k)
Principal AI Innovation Lead
Principal AI Innovation Lead

MultiPlan • McLean (VA)

On-site
USD 180,000 - 230,000
Health insurance
401k
Bonus opportunity
AI Quality Engineer
AI Quality Engineer

CitiusTech • Chicago (IL)

On-site
USD 140,000 - 195,000
Senior AI Agent QA Engineer for CI/CD & MLOps
Senior AI Agent QA Engineer for CI/CD & MLOps

MultiPlan • McLean (VA)

On-site
USD 140,000 - 160,000
Health insurance
401k
Bonus opportunity
AI Innovation, Principal
AI Innovation, Principal

Claritev • United States

On-site
USD 180,000 - 230,000
Health insurance
401k plan
Bonus opportunity
Director of AI, Provider Network
Director of AI, Provider Network

Claritev • United States

On-site
USD 210,000 - 230,000
Health insurance
401(k) with match
Employee stock purchase plan
+1