Member of Technical Staff - LLM Evals

Enclosure

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Enclosure in San Francisco is seeking an exceptional LLM evals researcher or engineer to own a significant portion of our LLM evals research as an IC. The role emphasizes exploration, curiosity, and practical impact in frontier AI security and model evaluation.

You will contribute to building robust evaluation environments and methods, collaborating with security and infrastructure teams to advance state-of-the-art AI safety and reliability.

Qualifications

  • Hands-on experience designing and running evaluations for frontier AI systems.
  • Experience building realistic evaluation environments, tasks, benchmarks, graders, or harnesses.
  • Strong experimental judgment in designing rigorous evaluations.
  • Experience evaluating tool-using agents, coding systems, and long-horizon tasks.
  • Strong ML and software fundamentals; RL or model training a strong plus.

Skills

Evaluation design
Experimentation
RL / model training
Software engineering fundamentals
Agent evals

Job description

About Enclosure

Enclosure is the world's first superintelligence security lab focused on preventing the sabotage, escape, and theft of ASI. We are building ASI-grade security for model weights, and nation-state-grade security for data centers.

We're backed by leaders across OpenAI, Anthropic, Google DeepMind, SpaceXAI, CoreWeave, NVIDIA, other frontier AI labs, and the national security community (including former NSA and intelligence community leadership). We are frontier AI, cybersecurity, quantum security, and national security researchers.

About the Role

Enclosure is hiring an exceptional LLM evals researcher or engineer. You'll own a significant part of our LLM evals research as an IC. We care less about degrees and more about a track record of exploration, curiosity, and grit.

Requirements

This role requires deep experience in multiple of the following areas:

  • Hands-on experience designing and running evaluations for frontier AI systems, ideally focused on cyber or other complex agentic capabilities

  • Experience building realistic evaluation environments, tasks, benchmarks, graders, or harnesses rather than only analyzing model outputs

  • Strong experimental judgment in designing rigorous evaluations

  • Experience evaluating tool-using agents, coding systems, and cyber and other long-horizon tasks

  • Strong ML and software engineering fundamentals, with experience in RL, post-training, or model training as a strong plus

What Sets You Apart

We're especially excited if you:

  • Demonstrate strong engineering judgment in the systems you've built

  • Get excited by hard problems and the challenge of figuring them out

  • Are highly adaptive and comfortable with how quickly the frontier shifts

  • Have evidence of self-directed technical work, such as open-source contributions, technical writing, tools, packages, or public projects with strong systems thinking

  • Communicate with clarity and directness

  • Bring a humble, low-ego attitude and care deeply about helping the team win

Why Join

Enclosure is securing the path to superintelligence.

What we offer:

  • Competitive compensation package

  • Ownership over research direction

  • Close collaboration with frontier AI security and infrastructure teams

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - AI Researcher
Member of Technical Staff - AI Researcher

Enclosure • San Francisco (CA)

On-site
USD 180,000 - 240,000
Ownership over research direction
Close collaboration with frontier AI安全
Member of Technical Staff - Cyber Researcher
Member of Technical Staff - Cyber Researcher

Enclosure • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Researcher, Reinforcement Learning
Member of Technical Staff - AI Researcher, Reinforcement Learning

Enclosure • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive compensation package
Ownership over research direction
Close collaboration with frontier AI"
Policy & Government Relations Lead
Policy & Government Relations Lead

Enclosure • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation package
Ownership over research direction
Close collaboration with frontier AI
Member of Technical Staff - Blade Runner
Member of Technical Staff - Blade Runner

Enclosure • San Francisco (CA)

On-site
USD 230,000 - 360,000
Competitive compensation
Ownership over research direction
Close collaboration with frontier AI"
Member of Technical Staff - AI Agents & Swarms
Member of Technical Staff - AI Agents & Swarms

Enclosure • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation package
Ownership over research direction
Close collaboration with frontier AI &
+1
Frontier LLM Evaluation Engineer - IC, Impact-Driven
Frontier LLM Evaluation Engineer - IC, Impact-Driven

Enclosure • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff - GPU Security
Member of Technical Staff - GPU Security

Enclosure • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation package
Ownership over research direction
Close collaboration with frontier AI"
Operations Generalist
Operations Generalist

Enclosure • San Francisco (CA)

On-site
USD 110,000 - 150,000
Competitive compensation package
Ownership over research direction
Close collaboration with frontier AI,
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity