Evals Engineer, Offensive Cyber

Zealot Labs

New York (NY)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Zealot Labs is seeking an Evals Engineer to own the truth layer for offensive cyber. You will shape the benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine what we trust, ship, and build next.

You’ll evaluate AI performance across vulnerability research, exploit development, CTF- and AIxCC-style challenges, and autonomous operations to answer whether non-deterministic agents can perform real offensive security work under realistic

Qualifications

  • Experience building LLM, model, or agent evaluation systems, with an understanding of contamination, benchmark overfit, grader drift, prompt sensitivity, brittle scoring, reward hacking, and false confidence.
  • Strong systems engineering ability; you can build reliable, reproducible evaluation infrastructure at scale.
  • A genuine security background. Hands‑on offensive experience in CTFs, vulnerability research, exploit development, reverse engineering, or AIxCC-style environments is a major advantage.
  • The range to read a heap‑corruption writeup and build the harness that determines whether a model actually landed the bug.
  • Rigor about measurement. These numbers inform real capability assessments and product and research decisions.

Responsibilities

  • Build CTF- and AIxCC-style task suites that reflect real targets and conditions, not toy problems.
  • Design eval methodology for non-deterministic agents, including variance across runs, sampling strategies, confidence thresholds, and statistically sound claims of capability.
  • Instrument full agent trajectories: tool calls, intermediate state, decision points, failed paths, partial progress, and final outcomes.
  • Build graders robust to reward hacking, including trustworthy LLM-as-judge pipelines and scoring for multi-stage exploitation chains.
  • Translate results into metrics and dashboards the research team uses to prioritize, and that hold up to scrutiny.

Skills

Evaluation systems
Systems engineering
Offensive security
CTF experience

Job description

Evals Engineer, Offensive Cyber

Our eval stack is the company’s truth layer. It tells us whether our AI systems can actually find vulnerabilities, reason through exploit chains, operate autonomously, and improve in ways that are real rather than cosmetic. If our evals are weak, we do not know what we have built. If they are strong, the entire research team moves faster with confidence.

We’re looking for an Evals Engineer to own that truth layer for offensive cyber: the benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine what we trust, what we ship, and what we build next.

You’ll evaluate AI performance across vulnerability research, exploit development, CTF- and AIxCC-style challenges, and autonomous operations. The core question: can non-deterministic agents perform real offensive security work under realistic conditions?

What You’ll Do
  • Build CTF- and AIxCC-style task suites that reflect real targets and conditions, not toy problems.
  • Design eval methodology for non-deterministic agents, including variance across runs, sampling strategies, confidence thresholds, and statistically sound claims of capability.
  • Instrument full agent trajectories: tool calls, intermediate state, decision points, failed paths, partial progress, and final outcomes.
  • Build graders robust to reward hacking, including trustworthy LLM-as-judge pipelines and scoring for multi-stage exploitation chains.
  • Translate results into metrics and dashboards the research team uses to prioritize, and that hold up to scrutiny.
What We’re Looking For
  • Experience building LLM, model, or agent evaluation systems, with an understanding of contamination, benchmark overfit, grader drift, prompt sensitivity, brittle scoring, reward hacking, and false confidence.
  • Strong systems engineering ability; you can build reliable, reproducible evaluation infrastructure at scale.
  • A genuine security background. Hands‑on offensive experience in CTFs, vulnerability research, exploit development, reverse engineering, or AIxCC-style environments is a major advantage.
  • The range to read a heap‑corruption writeup and build the harness that determines whether a model actually landed the bug.
  • Rigor about measurement. These numbers inform real capability assessments and product and research decisions.
Why This Role Matters

This is a high‑lever​age seat at the center of research, product, and engineering. The job is to turn messy agent behavior into evidence the team can trust.

The quality of everything we build depends on getting measurement right.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Offensive Cyber Eval Architect
Offensive Cyber Eval Architect

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
Cyber Evaluations Engineer
Cyber Evaluations Engineer

Anthropic • United States

Remote
USD 150,000 - 230,000
Research Engineer, Evaluations
Research Engineer, Evaluations

General Analysis • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 230,000
MTS - Research (Cybersecurity)
MTS - Research (Cybersecurity)

Collinear AI • Sunnyvale (CA)

On-site
USD 120,000 - 180,000
MTS - Research (Cybersecurity)
MTS - Research (Cybersecurity)

Collinear AI, Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Cyber Evaluations Engineer
Cyber Evaluations Engineer

EngineersOfAI • San Francisco (CA), Northern (KY)

On-site
USD 300,000 - 405,000
Applied Engineer, Evaluations Role
Applied Engineer, Evaluations Role

Mercor • United States

Remote
USD 120,000 - 180,000
Research Engineer, Evals
Research Engineer, Evals

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+6
Evaluation Engineer (Consulting / Contract-to-Hire) Seattle/Remote
Evaluation Engineer (Consulting / Contract-to-Hire) Seattle/Remote

AI Ethics Network • Northern (KY)

Hybrid
USD 138,000 - 220,000
Evals Lead
Evals Lead

Aslan • Washington

On-site
USD 120,000 - 180,000