Evaluation Engineer _ REMOTE(Flexible to work on PST timezone)

Momento USA LLC

United States

Hybrid

USD 90,000 - 170,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Momento USA LLC is seeking an Evaluation Engineer for a long-term contract, offering remote work with flexibility to align to PST timezone. The role centers on ensuring AI systems meet a defined quality bar before launch and through production, with ongoing monitoring.

You will build eval harnesses, create gold datasets with clinical reviewers, and set pass thresholds. Expect monthly reporting of pass rates and incidents, plus production drift monitoring in collaboration with the pipeline team.

Qualifications

  • 4 years in ML/LLM evaluation, QA engineering for AI systems, or applied research engineering.
  • Hands-on with eval frameworks and LLM-as-judge patterns and their failure modes, statistical rigor on small samples.
  • Independent spine: the team shipping a thing does not set its own pass bar.

Responsibilities

  • Build the central eval harness and templates every pod uses; hub-and-spoke central standards, pod-written tests
  • Create golden datasets with business and clinical reviewers, including safety and medical-accuracy suites
  • Set and defend pass thresholds; audit pod evals; report eval pass rates and incidents monthly
  • Stand up production monitoring for drift and regression with the pipeline engineer

Skills

ML/LLM evaluation
QA engineering
Statistical analysis
Eval frameworks

Job description

Momento USA is a global technology consulting, talent acquisition, and creative development firm that addresses clients’ most pressing needs and challenges. We are currently looking for a Evaluation Engineer.

Evaluation Engineer
REMOTE(Flexible to work on PST timezone)
Long term contract

Creates the quality bar every AI system must pass — golden datasets, eval harnesses, pass thresholds — before launch and continuously in production. Defines what “good enough to ship” means and measures it monthly for the scoreboard.

What they’ll do
  • Build the central eval harness and templates every pod uses; hub-and-spoke central standards, pod-written tests
  • Create golden datasets with business and clinical reviewers, including safety and medical-accuracy suites
  • Set and defend pass thresholds; audit pod evals; report eval pass rates and incidents monthly
  • Stand up production monitoring for drift and regression with the pipeline engineer
Must have
  • 4 years in ML/LLM evaluation, QA engineering for AI systems, or applied research engineering
  • Hands-on with eval frameworks, LLM-as-judge patterns and their failure modes, statistical rigor on small samples
  • Independent spine: the team shipping a thing does not set its own pass bar
Nice to have
  • Healthcare or safety-critical evaluation experience; red-teaming background
Note:

Momento USA is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status, or disability status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Evaluation Engineer — AI Quality & Safety
Remote Evaluation Engineer — AI Quality & Safety

Momento USA LLC • United States

Hybrid
USD 90,000 - 170,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Senior Artificial Intelligence Evaluation Engineer
Senior Artificial Intelligence Evaluation Engineer

Description Ciklum • United States

On-site
USD 120,000 - 190,000
Senior Artificial Intelligence Evaluation Engineer
Senior Artificial Intelligence Evaluation Engineer

Ciklum • United States

On-site
USD 120,000 - 170,000
Evals Lead
Evals Lead

Fluency Digital, Inc. • New York (NY)

On-site
USD 120,000 - 150,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna Inc. • San Francisco (CA)

On-site
USD 180,000 - 270,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required