Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Zealot Labs is seeking an Evals Engineer to own the truth layer for offensive cyber. You will shape the benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine what we trust, ship, and build next.
You’ll evaluate AI performance across vulnerability research, exploit development, CTF- and AIxCC-style challenges, and autonomous operations to answer whether non-deterministic agents can perform real offensive security work under realistic
Evals Engineer, Offensive Cyber
Our eval stack is the company’s truth layer. It tells us whether our AI systems can actually find vulnerabilities, reason through exploit chains, operate autonomously, and improve in ways that are real rather than cosmetic. If our evals are weak, we do not know what we have built. If they are strong, the entire research team moves faster with confidence.
We’re looking for an Evals Engineer to own that truth layer for offensive cyber: the benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine what we trust, what we ship, and what we build next.
You’ll evaluate AI performance across vulnerability research, exploit development, CTF- and AIxCC-style challenges, and autonomous operations. The core question: can non-deterministic agents perform real offensive security work under realistic conditions?
This is a high‑leverage seat at the center of research, product, and engineering. The job is to turn messy agent behavior into evidence the team can trust.
The quality of everything we build depends on getting measurement right.