Staff ML Evaluation & Calibration Engineer

Servicenow

Mountain View (CA)

On-site

USD 180,000 - 260,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Moveworks' AI agents don't just generate text — they act, plan, and change state in enterprise systems. You will own the central problem of scoring agent trajectories across multi-step interactions, turning complex behavior into meaningful judgments that train and improve the system.

This role focuses on designing reliable evaluation signals, combining deterministic validators with probabilistic LLM judgments, and calibrating a process reward model that informs improvements while guarding

Qualifications

  • 8+ years in applied ML, data science, or ML‑adjacent engineering with shipped work.
  • Ability to turn subjective judgement into measurable signals.
  • Strong Python and production-grade software discipline.
  • Experience evaluating LLMs and models in real-world settings.
  • Comfort with ambiguity and startup pace.

Responsibilities

  • Design and calibrate judge rubric systems for agent evaluation.
  • Develop deterministic validators and LLM judges for fuzzy judgments.
  • Build confidence reporting in scores and workflow integration.
  • Collaborate with annotation teams to align human labels.
  • Prototype and fine-tune small judge models.
  • Explain deviations and maintain robust evaluation pipelines.

Skills

Applied ML
Data science
ML engineering
Python
Experimentation
Communication

Tools

Evaluation design
LLM as judge
Prompt engineering
Fine-tuning models
Reward modeling

Job description

Moveworks' AI agents don't just generate text — they act, plan, and change state in enterprise systems. You will own the central problem of scoring agent trajectories across multi-step interactions, turning complex behavior into meaningful judgments that train and improve the system.

This role focuses on designing reliable evaluation signals, combining deterministic validators with probabilistic LLM judgments, and calibrating a process reward model that informs improvements while guarding

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Evaluation Systems Engineer
Senior AI Evaluation Systems Engineer

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Senior ML Engineer: Agentic Evaluation & Reward Modeling
Senior ML Engineer: Agentic Evaluation & Reward Modeling

Servicenow • Santa Clara (CA)

Hybrid
USD 150,000 - 230,000
Staff ML Engineer — Build Agentic AI for Enterprise
Staff ML Engineer — Build Agentic AI for Enterprise

JobCubby • San Diego (CA)

On-site
USD 150,000 - 190,000
Remote Staff ML Engineer - Agentic AI for Enterprise
Remote Staff ML Engineer - Agentic AI for Enterprise

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 230,000
Remote ML Engineer: Agentic AI Harness & Quality
Remote ML Engineer: Agentic AI Harness & Quality

Moveworks • Mountain View (CA)

On-site
USD 140,000 - 217,000
Health plans
401(k) match
ESPP
+3
Tech Lead for Agent Evaluation Platform
Tech Lead for Agent Evaluation Platform

Servicenow • Mountain View (CA)

Hybrid
USD 180,000 - 260,000
Staff AI Systems Engineer - LLM, Evaluation, Equity
Staff AI Systems Engineer - LLM, Evaluation, Equity

Maven • San Jose (CA)

On-site
USD 180,000 - 240,000
Physical Health Benefits
Mental Health Benefits
Emotional Health Benefits
+5
Staff Software Engineer — Scale AI Evaluation & Orchestration
Staff Software Engineer — Scale AI Evaluation & Orchestration

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 240,000
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Lead ML Engineer - NLU & Generative AI
Lead ML Engineer - NLU & Generative AI

ServiceNow • Mountain View (CA)

On-site
USD 180,000 - 240,000