Lead, Build Agent Evaluation & Model Benchmarking

ServiceNow

Santa Clara (CA)

On-site

USD 166,500 - 291,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health plans
401(k) plan with company match
ESPP
Matching donations
Flexible time away plan
Family leave programs

Job summary

ServiceNow in Santa Clara, CA is seeking a leader for the Build Agent evaluation framework. You will own eval strategy, roadmap, telemetry, and cross-team quality standards, guiding an 8‑engineer team to deliver scalable, data‑driven improvements.

You will benchmark models, decide on model support, and communicate results to stakeholders. Strong experience with evaluation systems for LLMs and cross‑team accountability is required.

Qualifications

  • Direct experience building or leading evaluation systems for LLM-based products, including scoring methodology and benchmark design.
  • Ability to drive cross-team accountability and get issues fixed from other teams based on data.
  • Working understanding of coding agent architecture: tool use, agent loops, context management.
  • 6+ years’ experience with ServiceNow-relevant technologies and advanced coding skills (Java, C++, Ruby, Shell, JavaScript).
  • Experience critically evaluating foundation models and distinguishing capability gaps from tuning artifacts.
  • Ability to navigate ambiguous priorities with context, risk, and outcomes.
  • People management experience at scale (8+ engineers) or managing managers is a plus.

Responsibilities

  • Own eval strategy and roadmap: prompt set design, scoring, failure taxonomy, telemetry.
  • Coordinate across multiple teams to ensure quality standards and drive fixes.
  • Lead model benchmarking and make data-backed recommendations on support decisions.
  • Manage daily activities of an 8-engineer team, including staffing and mentoring.
  • Address cross-cutting problems where eval data, model capability, and architecture intersect.
  • Present eval results and recommendations to stakeholders, including executives.

Skills

Direct experience building evaluation
Cross-team accountability
Coding agent architecture
Java/C++/Ruby/JS expertise
Evaluating foundation models
Ambiguous-priority execution
People management (8+ engineers)

Job description

ServiceNow in Santa Clara, CA is seeking a leader for the Build Agent evaluation framework. You will own eval strategy, roadmap, telemetry, and cross-team quality standards, guiding an 8‑engineer team to deliver scalable, data‑driven improvements.

You will benchmark models, decide on model support, and communicate results to stakeholders. Strong experience with evaluation systems for LLMs and cross‑team accountability is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Build Agent Engineering Manager
AI Build Agent Engineering Manager

ServiceNow • Santa Clara (CA)

On-site
USD 166,500 - 291,400
Health plans
401(k) Plan with company match
ESPP
+3
AI Eval & Benchmarking Manager - 8-Engineer Team Lead
AI Eval & Benchmarking Manager - 8-Engineer Team Lead

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
Software Engineering Manager - Build Agent
Software Engineering Manager - Build Agent

ServiceNow • Santa Clara (CA)

On-site
USD 166,000 - 292,000
Health plans
401(k) plan with company match
ESPP
+3
Software Engineering Manager - Build Agent
Software Engineering Manager - Build Agent

ServiceNow • Santa Clara (CA)

On-site
USD 166,500 - 291,400
Health plans
401(k) Plan with company match
ESPP
+3
Senior AI Systems Engineer, Evaluation & Orchestration
Senior AI Systems Engineer, Evaluation & Orchestration

Servicenow • Mountain View (CA)

On-site
USD 161,000 - 274,000
Health plans
401(k) with company match
Employee stock purchase plan (ESPP)
+2
Software Engineering Manager - Build Agent
Software Engineering Manager - Build Agent

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
Head of Enterprise AI Platform & LLM Strategy
Head of Enterprise AI Platform & LLM Strategy

ServiceNow • Santa Clara (CA)

On-site
USD 251,000 - 385,000
Health plans
401(k) Plan with company match
Employee stock purchase plan (ESPP)
+1
Staff ML Engineer: Enterprise AI Agents, Production-ready
Staff ML Engineer: Enterprise AI Agents, Production-ready

ServiceNow • Mountain View (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Agents
Staff Engineer, Agents

LeanData • Santa Clara (CA)

On-site
USD 160,000 - 200,000
Senior AI Evaluation & GenAI Benchmarking Lead
Senior AI Evaluation & GenAI Benchmarking Lead

ServiceNow • Santa Clara (CA)

On-site
USD 201,000 - 353,000