AI Benchmarking & Automation Engineer

ServiceNow

San Diego (CA)

On-site

USD 150,000 - 190,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Magnit Global benefits

Job summary

ServiceNow in San Diego, CA, is seeking an engineer to help build and scale tooling that measures AI-powered software development tools. You will develop evaluation harnesses, automate benchmark runs, and ensure reproducible results across teams.

The role emphasizes measurement quality over automation, with collaboration across engineering and data teams to document methods and findings for technical and leadership audiences.

Qualifications

  • Strong software engineering background with automation experience.
  • Experience with Git and Docker.
  • Experience with APIs, CI/CD, and engineering workflows.
  • Familiarity with AI, LLM, or agent evaluation and benchmarking.
  • Ability to troubleshoot and analyze results.

Responsibilities

  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

Skills

Software engineering
Automation
Developer tooling
Git
Docker
APIs
CI/CD pipelines
AI/LLM evaluation
Troubleshooting

Tools

Docker
CI/CD tooling

Job description

ServiceNow in San Diego, CA, is seeking an engineer to help build and scale tooling that measures AI-powered software development tools. You will develop evaluation harnesses, automate benchmark runs, and ensure reproducible results across teams.

The role emphasizes measurement quality over automation, with collaboration across engineering and data teams to document methods and findings for technical and leadership audiences.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Eval & Benchmarking Manager - 8-Engineer Team Lead
AI Eval & Benchmarking Manager - 8-Engineer Team Lead

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
AI Build Agent Engineering Manager
AI Build Agent Engineering Manager

ServiceNow • Santa Clara (CA)

On-site
USD 166,500 - 291,400
Health plans
401(k) Plan with company match
ESPP
+3
Senior AI Evaluation & GenAI Benchmarking Lead
Senior AI Evaluation & GenAI Benchmarking Lead

ServiceNow • Santa Clara (CA)

On-site
USD 201,000 - 353,000
Software Engineer - AI Eval & Automation
Software Engineer - AI Eval & Automation

ServiceNow • San Diego (CA)

On-site
USD 150,000 - 190,000
Magnit Global benefits
AI Tool Evaluation Engineer: Benchmarking & Automation
AI Tool Evaluation Engineer: Benchmarking & Automation

Akraya, Inc. • San Diego (CA)

On-site
USD 90,000 - 101,000
Senior AI Systems Engineer, Evaluation & Orchestration
Senior AI Systems Engineer, Evaluation & Orchestration

Servicenow • Mountain View (CA)

On-site
USD 161,000 - 274,000
Health plans
401(k) with company match
Employee stock purchase plan (ESPP)
+2
Senior AI-Native Systems Engineer
Senior AI-Native Systems Engineer

Servicenow • San Francisco (CA)

On-site
USD 191,000 - 334,000
Health plans
Flexible spending accounts
401(k) with company match
+4
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
AI-Driven Software Engineer: Build Scalable, Reusable Code
AI-Driven Software Engineer: Build Scalable, Reusable Code

Servicenow • West Palm Beach (FL)

Hybrid
USD 100,000 - 150,000
Senior AI Enablement Architect
Senior AI Enablement Architect

Servicenow • Santa Clara (CA)

On-site
USD 114,000 - 199,000
Health plans
401(k) Plan with company match
ESPP
+3