Remote Evaluation Engineer — AI Quality & Safety

Momento USA LLC

United States

Hybrid

USD 90,000 - 170,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Momento USA LLC is seeking an Evaluation Engineer for a long-term contract, offering remote work with flexibility to align to PST timezone. The role centers on ensuring AI systems meet a defined quality bar before launch and through production, with ongoing monitoring.

You will build eval harnesses, create gold datasets with clinical reviewers, and set pass thresholds. Expect monthly reporting of pass rates and incidents, plus production drift monitoring in collaboration with the pipeline team.

Qualifications

  • 4 years in ML/LLM evaluation, QA engineering for AI systems, or applied research engineering.
  • Hands-on with eval frameworks and LLM-as-judge patterns and their failure modes, statistical rigor on small samples.
  • Independent spine: the team shipping a thing does not set its own pass bar.

Responsibilities

  • Build the central eval harness and templates every pod uses; hub-and-spoke central standards, pod-written tests
  • Create golden datasets with business and clinical reviewers, including safety and medical-accuracy suites
  • Set and defend pass thresholds; audit pod evals; report eval pass rates and incidents monthly
  • Stand up production monitoring for drift and regression with the pipeline engineer

Skills

ML/LLM evaluation
QA engineering
Statistical analysis
Eval frameworks

Job description

Momento USA LLC is seeking an Evaluation Engineer for a long-term contract, offering remote work with flexibility to align to PST timezone. The role centers on ensuring AI systems meet a defined quality bar before launch and through production, with ongoing monitoring.

You will build eval harnesses, create gold datasets with clinical reviewers, and set pass thresholds. Expect monthly reporting of pass rates and incidents, plus production drift monitoring in collaboration with the pipeline team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Evaluation Engineer _ REMOTE(Flexible to work on PST timezone)
Evaluation Engineer _ REMOTE(Flexible to work on PST timezone)

Momento USA LLC • United States

Hybrid
USD 90,000 - 170,000
Senior AI Evaluation Engineer: AI Safety & Quality Lead
Senior AI Evaluation Engineer: AI Safety & Quality Lead

Description Ciklum • United States

On-site
USD 120,000 - 190,000
AI Implementation Quality & Evaluation Lead (Remote)
AI Implementation Quality & Evaluation Lead (Remote)

Granicus • United States

On-site
USD 80,000 - 105,700
Flexible Time Off
Company-Wide Wellbeing Days
Work From Home Reimbursement
+4
Remote Engineering AI Evaluator and Quality Reviewer
Remote Engineering AI Evaluator and Quality Reviewer

Obsidian • New York (NY)

On-site
USD 55,000 - 110,000
Remote Mobile AI Evaluation Specialist
Remote Mobile AI Evaluation Specialist

AuraOne • United States

On-site
USD 34,000 - 69,000
Remote AI Quality & Testing Evaluation Specialist
Remote AI Quality & Testing Evaluation Specialist

AuraOne • United States

On-site
USD 55,104 - 96,432
Remote AI Evaluation Specialist – Engineering/Manufacturing Ops
Remote AI Evaluation Specialist – Engineering/Manufacturing Ops

AuraOne • United States

On-site
USD 110,000 - 165,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Director, AI Evaluation & Quality
Director, AI Evaluation & Quality

United States Digital Space LLC • United States

Remote
USD 180,000 - 280,000
Senior AI Quality & Evaluation Scientist — Remote
Senior AI Quality & Evaluation Scientist — Remote

Matcha • Northern (KY)

Hybrid
USD 180,000 - 230,000