Agent Evaluation Platform Tech Lead

ServiceNow

Mountain View (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Generous family leave
Matched donations
Annual learning stipends
Flexible PTO
Competitive retirement plan
Paid volunteer time

Job summary

Moveworks is hiring to build the evaluation layer for its AI agent platform. You will own the judgment framework, rubrics, and calibration against human labels for multi-tenant, stateful enterprise environments.

The role focuses on applied ML at scale, not pretraining, with an emphasis on reliable instrumentation and reproducible results. You’ll work on the runtime that runs multi-turn agent scenarios, establish observability using OpenTelemetry, and guide the end-to-end delivery of a project

Qualifications

  • 10+ years building production backend or infrastructure systems.
  • Strong in Python or Go, with experience in scalable backend design.
  • Ability to lead end-to-end delivery of complex projects.
  • Excellent communication and collaboration skills.
  • Comfort with ambiguity and solving novel problems.

Responsibilities

  • Own eval orchestration at scale and the runtime that executes multi-turn agent scenarios end-to-end.
  • Lead the development of a single, reliable harness for agent evaluation across teams.
  • Establish observability, tracing, and reliability standards for the platform.
  • Collaborate with AI/ML, data, and infrastructure teams to ship robust tooling.

Skills

Python
Go
Backend systems
Leadership
Communication
OpenTelemetry

Job description

  • Moveworks’ AI agents don’t just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better?
  • That signal is what this role owns. You’ll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing
  • This isn’t a pretraining role, and it isn’t a testing role. It’s applied ML at a point where the methodology genuinely isn’t settled: LLMs judging LLMs is an open research problem, and we’re working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments
  • We’re hiring across three areas. You’ll anchor on one and touch the others; which one is a conversation we have with you, not a slot we drop you into
  • Eval orchestration at scale
  • The runtime that executes multi-turn agent scenarios end-to-end — stand up the environment and user simulator, drive the useragentworld loop, collect transcripts, traces, and final state, run validators and scoring, tear down
  • Scheduling, retries, high-concurrency execution, and run isolation at production dataset sizes
  • Versioned specs, datasets, and reports, with run-to-run comparison as a first-class operation
  • Consolidating evals that run today as one-off workflows onto a single orchestration service — one source of truth, one place to schedule and retry
  • Establishing a reliability floor and an SLO for the harness itself
  • Getting to self-serve, so any team runs an eval without bespoke integration
  • Agent observability and tracing
  • Leading the move to OpenTelemetry-native observability for the agent platform, replacing the parallel per-service logging, correlation, and redaction mechanisms in use today
  • The span data model for agent trajectories — prompts, tool calls, plan updates, outcomes — so a trajectory is queryable, not reconstructed by hand from log files
  • Trace context propagation across async boundaries and sessions that stay alive for minutes or hours
  • Making full prompts and completions survive the pipeline intact, and keeping eval traffic from contaminating its own data
  • Fault attribution and cross-run diffing: which component actually broke, and what changed since the last green run
  • The debug surface support and harness engineers use, and the tracing contract with the team that builds the agent
  • Stateful simulation
  • The simulation environment itself: stateful fakes of the enterprise systems agents call — ITSM, HR, knowledge bases, inventory — backed by a real datastore that persists changes during a run, so a created ticket is visible to a later read
  • Per-run data injection and programmatic setup/teardown so every run is hermetic and repeatable
  • LLM-driven user simulators for open-ended personas, and scripted state-machine simulators for deterministic flows
  • Contract-testing mocks against real API schemas in CI, so simulation fidelity can’t quietly drift as vendor APIs change
  • Ahead of us: isolated sandbox environments reproducing the config, identity, search content, and permissions an agent actually reads — provisioned from an identical baseline and torn down every run
  • And across all three: laying the foundation for using eval signal to optimize the agent, not just measure it
Benefits
  • Generous family leave
  • Matched donations
  • Annual learning stipends
  • Flexible PTO
  • Competitive retirement plan
  • Paid volunteer time

Concurrent and async programming: Python asyncio, Go concurrency, structured cancellationData-intensive pipelines: high-volume ingest, schema evolution, sampling and retention trade-offsGRPC/protobuf service and interface designObservability internals as a builder, not just a user: OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace dataOrchestration and workflow runtimes: DAG execution, scheduling, retries, backfills, high-concurrency job systems (Temporal, Airflow, Argo, or something you built yourself)Distributed systems: idempotency, delivery guarantees, isolation, and — unusually central here — determinism and reproducibilityComfort making a non-deterministic system measurable. You don’t need an ML background — but you should find it interesting to turn fuzzy agent behavior into a signal engineers are willing to gate releases on10+ years building production backend or infrastructure systemsStrong in Python or Go (ideally both)Experience designing and operating systems that handle real traffic at scaleAbility to tech lead other engineers and the end to end delivery of a project. Good communication and soft skillsComfort with ambiguity; these are novel problems without textbook solutions

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Evals Lead
Evals Lead

Aslan • Washington

On-site
USD 120,000 - 180,000
Member of Technical Staff
Member of Technical Staff

Chakra Labs • New York (NY)

On-site
USD 100,000 - 140,000
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 240,000
Tech Lead, Agent Eval Platform
Tech Lead, Agent Eval Platform

Servicenow • Mountain View (CA)

On-site
USD 180,000 - 260,000
Senior Platform Engineer
Senior Platform Engineer

Rifa AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Healthcare
Stock options
Agent Engineer
Agent Engineer

Rifa AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Agent Systems Engineer
Agent Systems Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Flexible work
Travel stipend
Lunch stipend
+1
Member of Technical Staff: Agent Runtime
Member of Technical Staff: Agent Runtime

ego AI (YC W24) • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Agent Harness Engineer
Agent Harness Engineer

Axiom • San Francisco (CA)

Hybrid
USD 180,000 - 260,000