Offensive Cyber Eval Architect

Zealot Labs

New York (NY)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Zealot Labs is seeking an Evals Engineer to own the truth layer for offensive cyber. You will shape the benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine what we trust, ship, and build next.

You’ll evaluate AI performance across vulnerability research, exploit development, CTF- and AIxCC-style challenges, and autonomous operations to answer whether non-deterministic agents can perform real offensive security work under realistic

Qualifications

  • Experience building LLM, model, or agent evaluation systems, with an understanding of contamination, benchmark overfit, grader drift, prompt sensitivity, brittle scoring, reward hacking, and false confidence.
  • Strong systems engineering ability; you can build reliable, reproducible evaluation infrastructure at scale.
  • A genuine security background. Hands‑on offensive experience in CTFs, vulnerability research, exploit development, reverse engineering, or AIxCC-style environments is a major advantage.
  • The range to read a heap‑corruption writeup and build the harness that determines whether a model actually landed the bug.
  • Rigor about measurement. These numbers inform real capability assessments and product and research decisions.

Responsibilities

  • Build CTF- and AIxCC-style task suites that reflect real targets and conditions, not toy problems.
  • Design eval methodology for non-deterministic agents, including variance across runs, sampling strategies, confidence thresholds, and statistically sound claims of capability.
  • Instrument full agent trajectories: tool calls, intermediate state, decision points, failed paths, partial progress, and final outcomes.
  • Build graders robust to reward hacking, including trustworthy LLM-as-judge pipelines and scoring for multi-stage exploitation chains.
  • Translate results into metrics and dashboards the research team uses to prioritize, and that hold up to scrutiny.

Skills

Evaluation systems
Systems engineering
Offensive security
CTF experience

Job description

Zealot Labs is seeking an Evals Engineer to own the truth layer for offensive cyber. You will shape the benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine what we trust, ship, and build next.

You’ll evaluate AI performance across vulnerability research, exploit development, CTF- and AIxCC-style challenges, and autonomous operations to answer whether non-deterministic agents can perform real offensive security work under realistic

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Evals Engineer, Offensive Cyber
Evals Engineer, Offensive Cyber

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
Chief Engineering Leader, AI-Driven Cyber R&D
Chief Engineering Leader, AI-Driven Cyber R&D

Zealot Labs, Inc. • Northern (KY), New York (NY)

Hybrid
USD 230,000 - 260,000
Equity in company
Head of Engineering
Head of Engineering

Snatch UP Jobs • New York (NY)

On-site
USD 230,000 - 260,000
Equity
Cyber Evaluations Engineer
Cyber Evaluations Engineer

Anthropic • United States

Remote
USD 150,000 - 230,000
AI / Systems Engineer, Offensive Cybersecurity
AI / Systems Engineer, Offensive Cybersecurity

Zealotlabs • Washington

On-site
USD 160,000 - 250,000
AI Safety & Cyber Evaluations Engineer
AI Safety & Cyber Evaluations Engineer

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 300,000 - 405,000
Head of Engineering LEADERSHIP CYBER R&D VR EMBEDDED WINDOWS New York, NY / On Site →
Head of Engineering LEADERSHIP CYBER R&D VR EMBEDDED WINDOWS New York, NY / On Site →

Zealot Labs, Inc. • Northern (KY), New York (NY)

Hybrid
USD 230,000 - 260,000
Equity in company
AI Engineer: Offensive Security Agents
AI Engineer: Offensive Security Agents

Zealot • United States

On-site
USD 100,000 - 140,000
Competitive cash
Meaningful early equity
Solving hard problems
Head of Engineering LEADERSHIP CYBER R&D VR EMBEDDED WINDOWS Washington, DC / On Site →
Head of Engineering LEADERSHIP CYBER R&D VR EMBEDDED WINDOWS Washington, DC / On Site →

Zealot Labs, Inc. • Washington

On-site
USD 230,000 - 260,000
Cybersecurity Evaluations Engineer (AI Safety & Robustness)
Cybersecurity Evaluations Engineer (AI Safety & Robustness)

Anthropic • United States

Remote
USD 150,000 - 230,000