Evals Lead

Aslan

Washington (District of Columbia)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Aslan in Washington DC is seeking a role focused on evaluation and quality for autonomous agents. You will own daily platform operation, build regression datasets, and drive the release process for models and harnesses.

You bring hands-on experience with evals for shipped LLM or agent products, and you can defend proxy metrics when the real signal isn't available, ensuring robust evaluation in production.

Qualifications

  • Strong Python
  • An evals framework you've run in production: Inspect, Promptfoo, Braintrust, Arize, LangSmith or equivalent
  • Familiarity with agent harness internals: memory, context assembly, tool selection
  • Preference data, process supervision, or reward modeling experience
  • Enough statistics to tell a regression from noise on small samples

Responsibilities

  • Run the platform yourself every day and log what's wrong in enough detail that an engineer can reproduce it.
  • Sample agent plans, rate them, and turn ratings into a versioned dataset.
  • Diagnose whether a bad decision came from the model or from the harness handing it the wrong context, and route each to the right owner.
  • Write and run checks over each agent's full history, and over properties that only hold across the whole set of running agents.
  • Build a simulation harness so a week of agent operation runs overnight against a candidate build.
  • Define what counts as a regression between builds. Measure both whether agents last and whether they get anywhere.

Skills

Python
Evals framework in production
Agent harness internals
Preference data / supervision / reward
Statistics for regression

Job description

About Aslan

Our national security apparatus was designed for an adversary that congregated and operate in the physical world. The threats disrupting our way of life today have gone faceless, borderless, and beyond the reach of any human operator. Aslan builds the autonomous agents that reach them, unmask them, and stop them, before a life can be lost, a household can be bankrupted, innocence can be stolen, or our can be nation subverted.

In 12 months we've gone from founding to live operational tasking and pilots across multiple national security and law enforcement partners. We've proven it works. This role makes it permanent.

We’ve raised $20M to date from Khosla Ventures, XYZ, BoxGroup, 2048, Liquid 2, Precursor Ventures, and others.

Role

Aslan builds autonomous agents that run for weeks at a time across many external systems, plan their own next steps, and work toward an objective. Scoring them the usual way doesn't work. Individual outputs look fine. Results come back late or not at all. What we need to know is whether the agent chose well at the moment it chose, and nothing off the shelf measures that.

This is our first quality role. You'll operate the platform daily, build the tooling that catches what you catch by hand, and own the decision to ship. We deploy on-premise on a fixed cadence and can't push fixes after delivery.

Responsibilities
  • Run the platform yourself every day and log what's wrong in enough detail that an engineer can reproduce it.

  • Sample agent plans, get them rated on whether the agent made the right call, and turn the ratings into a versioned dataset. Engineering uses it as a regression suite. The training side uses it as labels.

  • Diagnose whether a bad decision came from the model or from the harness handing it the wrong context, and route each to the right owner.

  • Write and run checks over each agent's full history, and over properties that only hold across the whole set of running agents.

  • Build a simulation harness so a week of agent operation runs overnight against a candidate build.

  • Define what counts as a regression between builds. Measure both whether agents last and whether they get anywhere. A build that only makes them cautious is a failed build.

Who You Are
  • You've owned evals for a shipped LLM or agent product, and you can say what your suite caught and what it missed

  • You've built a labeled dataset that got used for training, not just for measurement

  • You've picked a proxy metric because the real signal wasn't available, and you can defend the choice

  • You know when LLM-as-judge works and when it doesn't

  • You'd rather write the check than file the ticket

What You Bring
  • Strong Python

  • An evals framework you've run in production: Inspect, Promptfoo, Braintrust, Arize, LangSmith or equivalent

  • Familiarity with agent harness internals: memory, context assembly, tool selection

  • Preference data, process supervision, or reward modeling experience

  • Enough statistics to tell a regression from noise on small samples

Bonus Points
  • Evaluating agents that run for days or weeks, where per-episode metrics don't work

  • Systems you can't roll back or reset

  • Graph stores, provenance, lineage tracking

  • Eligible for a US security clearance

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Agent Evaluation Platform Tech Lead
Agent Evaluation Platform Tech Lead

ServiceNow • Mountain View (CA)

On-site
USD 180,000 - 240,000
Generous family leave
Matched donations
Annual learning stipends
+3
AI Engineer
AI Engineer

Aslan • Washington

On-site
USD 150,000 - 210,000
Agent Harness Engineer
Agent Harness Engineer

Axiom • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Member of Technical Staff: Agent Runtime
Member of Technical Staff: Agent Runtime

ego AI (YC W24) • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Member of Technical Staff, AI Products
Member of Technical Staff, AI Products

Zaun • San Francisco (CA)

On-site
USD 150,000 - 210,000
Forward Deployed Engineer
Forward Deployed Engineer

Aslan • Washington

On-site
USD 120,000 - 190,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Engineer, Agent Builder (Remote).
AI Engineer, Agent Builder (Remote).

Catalyst Wayfare • United States

Remote
USD 120,000 - 170,000
Software Engineer, AI Evaluation
Software Engineer, AI Evaluation

Nuna • San Francisco (CA)

On-site
USD 150,000 - 230,000