LLM Evaluation Engineer – AI Agent Platforms

Servicenow

Santa Clara (CA)

On-site

USD 126,000 - 195,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health plans
401(k) Plan with company match
ESPP
Flexible time away
Family leave

Job summary

ServiceNow is seeking an expert in AI evaluation to advance its Build Agent team, the AI coding assistant for the platform. You will work on evaluation infrastructure, scoring, and failure analysis to ensure high-quality model performance across ServiceNow metadata types and workflows.

Responsibilities include large-scale eval orchestration and improving agent observability and tracing. A track record in LLM evaluation and model benchmarking is required; you will influence release quality and

Qualifications

  • Direct experience building evaluation systems for LLM-based products, including scoring methods and benchmarks.
  • Understanding of coding agent architecture: tools, loops, context management, scaffolding tradeoffs.
  • Experience evaluating foundation models and distinguishing real capability from tuning artifacts.
  • Ability to handle ambiguous priorities with incomplete eval data.

Responsibilities

  • Eval orchestration at scale
  • Agent observability and tracing

Skills

LLM evaluation systems
Agent architecture understanding
Foundation models evaluation
Ambiguity handling

Job description

ServiceNow is seeking an expert in AI evaluation to advance its Build Agent team, the AI coding assistant for the platform. You will work on evaluation infrastructure, scoring, and failure analysis to ensure high-quality model performance across ServiceNow metadata types and workflows.

Responsibilities include large-scale eval orchestration and improving agent observability and tracing. A track record in LLM evaluation and model benchmarking is required; you will influence release quality and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: Agentic Evaluation & Reward Modeling
Senior ML Engineer: Agentic Evaluation & Reward Modeling

Servicenow • Santa Clara (CA)

Hybrid
USD 150,000 - 230,000
AI Build Agent Engineering Manager
AI Build Agent Engineering Manager

ServiceNow • Santa Clara (CA)

On-site
USD 166,500 - 291,400
Health plans
401(k) Plan with company match
ESPP
+3
Tech Lead for Agent Evaluation Platform
Tech Lead for Agent Evaluation Platform

Servicenow • Mountain View (CA)

Hybrid
USD 180,000 - 260,000
AI Eval & Benchmarking Manager - 8-Engineer Team Lead
AI Eval & Benchmarking Manager - 8-Engineer Team Lead

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
LLM Agent Evaluation Engineer — Build Intelligent Agents
LLM Agent Evaluation Engineer — Build Intelligent Agents

Bytedance • San Jose (CA)

On-site
USD 154,000 - 256,000
Staff AI Systems Engineer - LLM, Evaluation, Equity
Staff AI Systems Engineer - LLM, Evaluation, Equity

Maven • San Jose (CA)

On-site
USD 180,000 - 240,000
Physical Health Benefits
Mental Health Benefits
Emotional Health Benefits
+5
Lead, Build Agent Evaluation & Model Benchmarking
Lead, Build Agent Evaluation & Model Benchmarking

ServiceNow • Santa Clara (CA)

On-site
USD 166,000 - 292,000
Health plans
401(k) plan with company match
ESPP
+3
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000