A complete application in a minute — tailored resume and cover letter, ready to send.
Enclosure in San Francisco is seeking an exceptional LLM evals researcher or engineer to own a significant portion of our LLM evals research as an IC. The role emphasizes exploration, curiosity, and practical impact in frontier AI security and model evaluation.
You will contribute to building robust evaluation environments and methods, collaborating with security and infrastructure teams to advance state-of-the-art AI safety and reliability.
Enclosure is the world's first superintelligence security lab focused on preventing the sabotage, escape, and theft of ASI. We are building ASI-grade security for model weights, and nation-state-grade security for data centers.
We're backed by leaders across OpenAI, Anthropic, Google DeepMind, SpaceXAI, CoreWeave, NVIDIA, other frontier AI labs, and the national security community (including former NSA and intelligence community leadership). We are frontier AI, cybersecurity, quantum security, and national security researchers.
Enclosure is hiring an exceptional LLM evals researcher or engineer. You'll own a significant part of our LLM evals research as an IC. We care less about degrees and more about a track record of exploration, curiosity, and grit.
This role requires deep experience in multiple of the following areas:
Hands-on experience designing and running evaluations for frontier AI systems, ideally focused on cyber or other complex agentic capabilities
Experience building realistic evaluation environments, tasks, benchmarks, graders, or harnesses rather than only analyzing model outputs
Strong experimental judgment in designing rigorous evaluations
Experience evaluating tool-using agents, coding systems, and cyber and other long-horizon tasks
Strong ML and software engineering fundamentals, with experience in RL, post-training, or model training as a strong plus
We're especially excited if you:
Demonstrate strong engineering judgment in the systems you've built
Get excited by hard problems and the challenge of figuring them out
Are highly adaptive and comfortable with how quickly the frontier shifts
Have evidence of self-directed technical work, such as open-source contributions, technical writing, tools, packages, or public projects with strong systems thinking
Communicate with clarity and directness
Bring a humble, low-ego attitude and care deeply about helping the team win
Enclosure is securing the path to superintelligence.
What we offer:
Competitive compensation package
Ownership over research direction
Close collaboration with frontier AI security and infrastructure teams