Senior ML Engineer: Agentic Evaluation & Reward Modeling

Servicenow

Santa Clara (CA)

Hybrid

USD 150,000 - 230,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ServiceNow seeks an experienced ML-focused engineer to own the evaluation framework for AI agents. You will design judges, calibration rubrics, and confidence-aware scoring, working across deterministic validators and LLM judgments to improve agent performance in enterprise workflows.

The role emphasizes building a production-grade evaluation platform, collaborating with annotation teams, and advancing self-learning for agent harnesses with rigorous measurement and ownership at startup pace.

Qualifications

  • 5+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used.
  • Experience turning subjective human judgement into a measurement that holds up — one that other models can act on.
  • Strong Python, and the discipline to ship production-grade code rather than notebooks.

Responsibilities

  • Design and calibrate judges and evaluation rubrics for agent trajectories.
  • Build and maintain an evaluation platform with deterministic validators and LLM-based judgments.
  • Develop a calibration loop with human-labeled trajectories and collaborate with annotation teams.

Skills

Applied ML experience
Python programming
Measurement/ evaluation design
LLM experience
Technical communication

Job description

ServiceNow seeks an experienced ML-focused engineer to own the evaluation framework for AI agents. You will design judges, calibration rubrics, and confidence-aware scoring, working across deterministic validators and LLM judgments to improve agent performance in enterprise workflows.

The role emphasizes building a production-grade evaluation platform, collaborating with annotation teams, and advancing self-learning for agent harnesses with rigorous measurement and ownership at startup pace.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluation Engineer – AI Agent Platforms
LLM Evaluation Engineer – AI Agent Platforms

Servicenow • Santa Clara (CA)

On-site
USD 126,000 - 195,000
Equity
Health plans
401(k) Plan with company match
+3
Tech Lead for Agent Evaluation Platform
Tech Lead for Agent Evaluation Platform

Servicenow • Mountain View (CA)

Hybrid
USD 180,000 - 260,000
Staff ML Evaluation & Calibration Engineer
Staff ML Evaluation & Calibration Engineer

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 260,000
AI Eval & Benchmarking Manager - 8-Engineer Team Lead
AI Eval & Benchmarking Manager - 8-Engineer Team Lead

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
Senior Machine Learning Engineer, Agent Eval Platform
Senior Machine Learning Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

Hybrid
USD 150,000 - 230,000
AI Build Agent Engineering Manager
AI Build Agent Engineering Manager

ServiceNow • Santa Clara (CA)

On-site
USD 166,500 - 291,400
Health plans
401(k) Plan with company match
ESPP
+3
Senior AI Systems Engineer, Evaluation & Orchestration
Senior AI Systems Engineer, Evaluation & Orchestration

Servicenow • Mountain View (CA)

On-site
USD 161,000 - 274,000
Health plans
401(k) with company match
Employee stock purchase plan (ESPP)
+2
Staff Machine Learning Engineer, Agent Eval Platform
Staff Machine Learning Engineer, Agent Eval Platform

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 260,000
Staff ML Engineer, Agentic AI: Lead with Generative Tech
Staff ML Engineer, Agentic AI: Lead with Generative Tech

ServiceNow • Mountain View (CA)

On-site
USD 130,000 - 160,000
Senior AI Evaluation & GenAI Benchmarking Lead
Senior AI Evaluation & GenAI Benchmarking Lead

ServiceNow • Santa Clara (CA)

On-site
USD 201,000 - 353,000