Senior AI Agent Quality Engineer

Claritev Corporation

United States

On-site

USD 140,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with company match
Bonus opportunity

Job summary

Claritev seeks a Senior AI Operations Engineer to gatekeep agentic AI systems powering healthcare products. You will design and run tests, benchmarks, and evaluation pipelines to catch regressions, hallucinations, and unsafe actions before release.

Collaborate with AI/ML teams and product stakeholders to build golden datasets, automate tests in CI/CD, triage agent failures, and translate results into quality insights while ensuring PHI/PII handling and HIPAA compliance.

Qualifications

  • Bachelor's degree in CS, engineering, data science, or related field.
  • 5+ years in software QA/test automation or ML engineering.
  • 2+ years with production ML/AI systems.
  • 1+ years with generative AI, LLMs, or agentic AI.
  • Track record building automation improving software/model quality.

Responsibilities

  • Design, build, and maintain automated test suites and evaluation pipelines for agentic AI systems.
  • Develop golden datasets and simulation environments reflecting healthcare workflows.
  • Automate agent benchmarking and CI/CD quality gates.
  • Measure and report agent quality metrics and safety compliance.
  • Implement LLM-as-judge and rubric-based scoring; verify with human review.
  • Triage, reproduce, and root-cause agent failures; file defect reports.
  • Perform adversarial testing and guardrail verification.
  • Monitor production agent behavior and feed findings into offline tests.
  • Contribute to evaluation dashboards, runbooks, and quality docs.
  • Ensure PHI/PII handling and HIPAA compliance; support data governance.

Skills

Python
Test automation
QA / SDET
ML engineering
CI/CD
Observability
Debugging
Data analysis
Communication

Education

Bachelor's degree in Computer Science or related field
Master's degree a plus

Tools

pytest
LangSmith
Langfuse
Arize Phoenix
Ragas
DeepEval
promptfoo
LangGraph
LangChain
AutoGen
CrewAI
Kubernetes
Terraform

Job description

Job Description

We are seeking a Senior AI Operations Engineer to serve as the quality gatekeeper for the agentic AI systems powering Claritev's next generation of healthcare products. This is a hands‑on Agentic QA role for an engineer who is passionate about answering a deceptively hard question: how do you prove an AI agent works? You will design and run the test suites, benchmarks, and evaluation pipelines that validate agent behavior before and after release — catching regressions, hallucinations, broken tool calls, and unsafe actions before they reach production healthcare workflows. You will work closely with our Principal Agentic AI Operations Engineer, AI engineering teams, and Product stakeholders to build golden datasets, automate agent testing in CI/CD, triage and reproduce agent failures, and turn evaluation results into actionable quality insights. This role is an excellent path for a strong QA/test automation, MLOps, or backend engineer looking to specialize in the fast‑growing discipline of AI agent quality and evaluation.

Job Roles And Responsibilities
  • Design, build, and maintain automated test suites and evaluation pipelines for agentic AI systems, covering single-turn, multi-turn, and end-to-end workflow scenarios.
  • Develop and curate golden datasets, test scenarios, and simulation environments that reflect real healthcare workflows and edge cases.
  • Execute and automate agent benchmarking, including regression testing across model, prompt, and tool changes, with clear pass/fail quality gates in CI/CD.
  • Measure and report agent quality metrics: task completion, tool-call accuracy, grounding/hallucination rates, response quality, latency, cost per task, and safety compliance.
  • Implement LLM-as-judge and rubric-based scoring approaches, and validate automated scores against human review.
  • Triage, reproduce, and root-cause agent failures by analyzing traces of LLM calls, tool invocations, retrieval steps, and orchestration paths; file actionable defect reports.
  • Perform adversarial and negative testing, including prompt-injection probes, malformed inputs, boundary conditions, and guardrail verification.
  • Monitor production agent behavior, investigate quality incidents and drift, and feed findings back into offline test coverage.
  • Contribute to evaluation dashboards, runbooks, and quality documentation used by engineering and product teams.
  • Verify PHI/PII handling, auditability, and human-in-the-loop controls behave as designed, supporting HIPAA and data-governance compliance.
  • Collaborate with AI engineers and data scientists to make agents more testable, observable, and reliable by design.
Job Requirements
Education
  • Bachelor's degree in Computer Science, Engineering, Data Science, a quantitative discipline, or a related field required.
  • Master's degree a plus.
Experience
  • 5+ years of hands‑on experience in software engineering, QA/test automation engineering, ML engineering, or a related technical discipline.
  • 2+ years of experience testing, operating, or supporting production ML or AI systems.
  • 1+ years of hands‑on experience with generative AI, LLMs, RAG, and/or agentic AI systems (professional projects preferred; substantial personal or open‑source work considered).
  • Track record of building test automation or evaluation capabilities that measurably improved software or model quality.
Technical Skills
  • Strong Python skills, including test frameworks (e.g., pytest) and building automation/tooling; comfort working with APIs and distributed systems.
  • Hands‑on experience with LLM/agent evaluation concepts: eval datasets, LLM-as-judge, rubric-based scoring, and regression detection.
  • Familiarity with evaluation and observability tooling such as LangSmith, Langfuse, Arize Phoenix, Ragas, DeepEval, promptfoo, or equivalent.
  • Working knowledge of agentic AI frameworks and patterns (e.g., LangGraph, LangChain, AutoGen, CrewAI), including tool use, orchestration, and guardrails.
  • Understanding of RAG pipelines, embeddings, and vector search, and how retrieval quality affects agent behavior.
  • Experience with CI/CD pipelines and integrating automated tests and quality gates into them.
  • Familiarity with cloud environments, containers, and log/trace analysis for debugging distributed systems.
  • Solid grasp of statistics fundamentals for interpreting evaluation results (sampling, variance, significance).
Other Skills
  • Strong analytical and debugging skills, with high attention to detail and a healthy skepticism toward “it works on my machine.”
  • Clear written communication, especially in defect reports, test plans, and quality summaries.
  • Ability to operate effectively in a fast‑moving, cross‑functional environment.
Preferred Qualifications
  • Experience with Oracle Cloud Infrastructure (OCI), including its generative AI, data science, database, and observability capabilities.
  • Experience in healthcare, health technology, insurance, claims, payment integrity, or other regulated industries.
  • Experience testing systems that process sensitive data, including PHI or PII, in HIPAA‑regulated environments.
  • Exposure to public agent/LLM benchmarks (e.g., SWE‑bench, GAIA, AgentBench, tau‑bench) or red‑teaming and safety evaluation of LLM systems.
  • Experience with Kubernetes, infrastructure‑as‑code (e.g., Terraform), or SRE practices.
  • Prior QA/SDET background applied to ML or data‑intensive systems.
Compensation

The salary range for this position is $140K-160K/year. Specific offers take into account a candidate’s education, experience and skills, as well as the candidate’s work location and internal equity. This position is also eligible for health insurance, 401k and bonus opportunity.

About Us

What Guides Us

At Claritev, innovation, agility, and a focus on results drive our success. We embrace bold thinking, work as one team, take ownership, and strive for excellence in everything we do — creating meaningful impact for our clients, communities, and each other.

Why Claritev?

Healthcare is complex. We help make it clearer. At Claritev, you'll do work that matters. Together, we're helping make healthcare more transparent and affordable for all through the power of data, technology, and expertise. We offer meaningful opportunities to grow your career, collaborate with talented colleagues, and make an impact on the clients and communities we serve. If you're looking for purpose, growth, and a team that succeeds together, you'll find it here.

About The Team

My Total Value – Claritev’s Global Total Rewards Philosophy – My Total Value – Is Grounded In Investing In Our Associates And Rewarding Performance. This Includes

  • Competitive compensation and incentive opportunities (where eligible)
  • Medical, dental, and vision coverage
  • Life and disability coverage
  • 401(k) with company match
  • Employee stock purchase plan
  • Flexible spending accounts and health savings accounts
  • Paid time off and company holidays
  • Paid parental leave
  • Mental health resources
  • Tuition reimbursementFinancial planning resourcesProfessional development opportunities
Our Commitment to Inclusion

At Claritev, we believe diverse perspectives strengthen our teams and lead to better outcomes. We are committed to fostering an inclusive workplace where every associate feels respected, valued, and empowered to succeed. Claritev is an Equal Opportunity Employer and complies with all applicable laws and regulations. Qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Agent Quality Engineer
Senior AI Agent Quality Engineer

MultiPlan • McLean (VA)

On-site
USD 140,000 - 160,000
Health insurance
401k
Bonus opportunity
Principal AI Engineer - Innovation
Principal AI Engineer - Innovation

Claritev Corporation • United States

On-site
USD 190,000 - 210,000
Health insurance
401(k) plan
Employee stock purchase plan
+1
Principal Agentic AI Operations Engineer
Principal Agentic AI Operations Engineer

MultiPlan • McLean (VA)

On-site
USD 165,000 - 185,000
Health insurance
401k
Bonus opportunity
Senior Applied AI Scientist
Senior Applied AI Scientist

Claritev • United States

On-site
USD 120,000 - 150,000
Medical, dental, and vision coverage
401(k) with company match
Paid time off
Principal Engineer/Architect, AI Innovation
Principal Engineer/Architect, AI Innovation

Claritev • United States

On-site
USD 180,000 - 230,000
Health insurance
Bonus opportunity
401(k)
Principal Applied AI Engineer, Agentic AI
Principal Applied AI Engineer, Agentic AI

Claritev • United States

On-site
USD 190,000 - 210,000
Director of AI, Provider Network
Director of AI, Provider Network

Claritev • United States

On-site
USD 210,000 - 230,000
Health insurance
401(k) with match
Employee stock purchase plan
+1
Lead Applied AI Engineer, Agentic AI
Lead Applied AI Engineer, Agentic AI

MultiPlan • McLean (VA)

On-site
USD 140,000 - 170,000
Health insurance
401k
Bonus opportunity
Principal AI Innovation Lead
Principal AI Innovation Lead

Claritev • United States

On-site
USD 180,000 - 230,000
Medical, dental, and vision coverage
401(k) with company match
Paid time off and holidays
+1
Principal Applied AI Engineer, Agentic AI
Principal Applied AI Engineer, Agentic AI

MultiPlan • McLean (VA)

On-site
USD 190,000 - 210,000
Health insurance
401(k) retirement plan
Performance bonus