Staff Software Engineer, Agent Eval Platform

Servicenow

Santa Clara (CA)

On-site

USD 180,000 - 320,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Moveworks is seeking an experienced backend/infrastructure engineer to design the judgment layer for our agent evaluation platform. You will define rubrics, calibrate labels, and shape the signal used to improve agents across multi-step trajectories.

This role focuses on applied ML at scale with stateful enterprise simulations, robust observability, and reliable backends. Strong Python or Go skills and comfort with ambiguity are essential.

Qualifications

  • 8+ years building production backend or infrastructure systems.
  • Strong in Python or Go (ideally both).
  • Experience designing and operating systems that handle real traffic at scale.
  • Comfort making a non-deterministic system measurable; ML background not required but helpful.
  • Comfort with ambiguity and evolving requirements.

Responsibilities

  • Own the evaluation platform's rubrics, judges, and calibration against human labels.
  • Build scalable orchestration for multi-turn agent scenarios and validation workflows.
  • Lead observability efforts, tracing, and data collection for agent trajectories.
  • Develop stateful simulation environments to reproduce enterprise systems and data.
  • Contribute to reliability, reproducibility, and run-to-run comparisons across experiments.

Skills

Distributed systems
Orchestration/workflows
Observability internals
Async programming
Data pipelines
gRPC/protobuf design

Tools

OpenTelemetry
Temporal/Airflow

Job description

The Role

Moveworks' AI agents don't just generate text - they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did - across a multi-step trajectory through a world it changed - precisely enough that the score can teach it to do better?

That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card - a judge good enough to grade a trajectory is a judge good enough to train against . The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.

This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

What you get to do in this role:

We're hiring across three areas. You'll anchor on one and touch the others; which one is a conversation we have with you, not a slot we drop you into.

Eval orchestration at scale
  • The runtime that executes multi-turn agent scenarios end-to-end - stand up the environment and user simulator, drive the useragentworld loop, collect transcripts, traces, and final state, run validators and scoring, tear down
  • Scheduling, retries, high-concurrency execution, and run isolation at production dataset sizes
  • Versioned specs, datasets, and reports, with run-to-run comparison as a first-class operation
  • Consolidating evals that run today as one-off workflows onto a single orchestration service - one source of truth, one place to schedule and retry
  • Establishing a reliability floor and an SLO for the harness itself
  • Getting to self-serve, so any team runs an eval without bespoke integration
Agent observability and tracing
  • Leading the move to OpenTelemetry-native observability for the agent platform, replacing the parallel per-service logging, correlation, and redaction mechanisms in use today
  • The span data model for agent trajectories - prompts, tool calls, plan updates, outcomes - so a trajectory is queryable , not reconstructed by hand from log files
  • Trace context propagation across async boundaries and sessions that stay alive for minutes or hours
  • Making full prompts and completions survive the pipeline intact, and keeping eval traffic from contaminating its own data
  • Fault attribution and cross-run diffing: which component actually broke, and what changed since the last green run
  • The debug surface support and harness engineers use, and the tracing contract with the team that builds the agent
Stateful simulation
  • The simulation environment itself: stateful fakes of the enterprise systems agents call - ITSM, HR, knowledge bases, inventory - backed by a real datastore that persists changes during a run, so a created ticket is visible to a later read
  • Per-run data injection and programmatic setup/teardown so every run is hermetic and repeatable
  • LLM-driven user simulators for open-ended personas, and scripted state-machine simulators for deterministic flows
  • Contract-testing mocks against real API schemas in CI, so simulation fidelity can't quietly drift as vendor APIs change
  • Ahead of us: isolated sandbox environments reproducing the config, identity, search content, and permissions an agent actually reads - provisioned from an identical baseline and torn down every run

And across all three: laying the foundation for using eval signal to optimize the agent, not just measure it.

To be successful in this role you have:
Experience in at least 3 of these:
  • Distributed systems: idempotency, delivery guarantees, isolation, and - unusually central here - determinism and reproducibility
  • Orchestration and workflow runtimes: DAG execution, scheduling, retries, backfills, high-concurrency job systems (Temporal, Airflow, Argo, or something you built yourself)
  • Observability internals as a builder , not just a user: OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace data
  • Concurrent and async programming: Python asyncio, Go concurrency, structured cancellation
  • Data-intensive pipelines: high-volume ingest, schema evolution, sampling and retention trade-offs
  • gRPC/protobuf service and interface design
Required:
  • 8+ years building production backend or infrastructure systems
  • Strong in Python or Go (ideally both)
  • Experience designing and operating systems that handle real traffic at scale
  • Comfort making a non-deterministic system measurable. You don't need an ML background - but you should find it interesting to turn fuzzy agent behavior into a signal engineers are willing to gate releases on
  • Comfort with ambiguity
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

Chakra Labs • New York (NY)

On-site
USD 100,000 - 140,000
Member of Technical Staff: Agent Runtime
Member of Technical Staff: Agent Runtime

ego AI (YC W24) • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Moveworks • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Agent Harness Engineer
Agent Harness Engineer

axiombio • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Agent Harness Engineer
Agent Harness Engineer

Axiom • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior Platform Engineer
Senior Platform Engineer

Rifa AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Healthcare
Stock options
Platform Engineer
Platform Engineer

ThirdLayer, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Senior Software Engineer, Agents
Senior Software Engineer, Agents

Jobtailor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Agent Engineer
Agent Engineer

Rifa AI • San Francisco (CA)

On-site
USD 120,000 - 180,000