Senior AI Quality Engineer (LLM Evaluation & Automation) 1754

Sokoni Kwetu Limited

United States

Remote

USD 95,000 - 140,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Sokoni Kwetu Limited is hiring for a remote role owning the eval harness and quality gate from day one. You will replace the old Evals Specialist model with a standing owner responsible for measurable agent quality.

You will design MVP-style evals, integrate them into CI, and establish thresholds with Product and the Tech Lead, paving the way for future adversarial testing while avoiding MVP overbuild.

Qualifications

  • Experience evaluating ML, LLM, or non-deterministic systems.
  • Strong test and benchmark design capability.
  • Comfort working with noisy metrics, thresholds, and probabilistic behavior.
  • Good scripting and automation skills.

Responsibilities

  • Build and maintain the MVP eval harness: golden tasks, exception tasks, scorecard metrics, and regression packs.
  • Wire evals into CI so quality regressions fail builds and releases.
  • Define and maintain release-gate thresholds with Product and the Tech Lead.
  • Lay the path for later adversarial and drift-testing expansion without overbuilding MVP scope.

Skills

ML evaluation
Benchmark design
Scripting
Automation
Probabilistic reasoning

Tools

CI/CD pipelines

Job description

This is a remote position.

Owns the eval harness and quality gate from the beginning. This role replaces the old late-stage “Evals Specialist” model with a standing owner for measurable agent quality.

Key Responsibilities
  • Build and maintain the MVP eval harness: golden tasks, exception tasks, scorecard metrics, and regression packs.
  • Wire evals into CI so quality regressions fail builds and releases.
  • Define and maintain release-gate thresholds with Product and the Tech Lead.
  • Lay the path for later adversarial and drift-testing expansion without overbuilding MVP scope.
Requisitos
Must-Have Qualifications
  • Experience evaluating ML, LLM, or non-deterministic systems.
  • Strong test and benchmark design capability.
  • Comfort working with noisy metrics, thresholds, and probabilistic behavior.
  • Good scripting and automation skills.
AI-First Expectations
  • Uses AI to generate candidate eval cases and failure hypotheses, but never confuses generated tests with validated quality.
  • Approaches AI quality as an operating system, not a QA afterthought.
What Success Looks Like in the First 90 Days
  • The first reference agent has a published scorecard and gated eval path.
  • Golden and exception tests run automatically.
  • The team can explain what "good enough to ship" means in measurable terms.

Originally posted on Himalayas

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Scientist — Agent Evaluations & Quality
Data Scientist — Agent Evaluations & Quality

Clera • United States

Remote
USD 120,000 - 190,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
Remote AI Quality Engineer: Eval Harness Lead
Remote AI Quality Engineer: Eval Harness Lead

Sokoni Kwetu Limited • United States

Remote
USD 95,000 - 140,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Town of Florida (NY), Northern (KY)

Hybrid
USD 120,000 - 150,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Agency • United States

Remote
USD 8,300 - 90,000
Senior Product Software Engineer - AI Quality Engineering
Senior Product Software Engineer - AI Quality Engineering

Wolters Kluwer - Financial Services Solutions • Cape Girardeau (MO)

On-site
USD 120,000 - 150,000