Research Engineer, Evaluations

General Analysis

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 230,000

Full time

38 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

General Analysis is seeking a Research Engineer to build the systems that define how frontier AI models are evaluated and secured. You will iterate on agent frameworks and build the scaffolding that lets models perform at their best, incorporating state-of-the-art tools for task delegation and complex multi-step tasks.

You will design and build evaluation environments, including challenges that measure a model's ability to evade detection.

Qualifications

  • Strong production programming skills and ability to move quickly from prototype to reliable infrastructure.
  • Experience building agentic systems, evaluation harnesses, or environment frameworks or eagerness to learn.
  • Security experience or interest in building security into AI systems.

Responsibilities

  • Build the systems that define how frontier AI models are evaluated and secured.
  • Iterate on agent frameworks and scaffolding for multi-step tasks.
  • Design and build evaluation environments including challenges to test model evasion capabilities.
  • Collaborate with red teamers, pentesters, and contracted security firms to ground experiments.
  • Own the pipeline end-to-end from API design to analysis and visualization of thousands of agent trajectories.
  • Publish results in blog posts and academic venues and deliver findings to customers.

Skills

Production coding
Agentic systems
Evaluation harnesses
Security experience
Analytical skills

Job description

General Analysis is a frontier security lab. We protect the world as AI systems grow more capable and more dangerous in the wrong hands. We build high-fidelity research infrastructure that evaluates and strengthens frontier models on critical cyberdefense tasks, and we bring that research directly to enterprises — securing the AI systems they deploy through adversarial testing and runtime protection.

About the role

As a Research Engineer, you will build the systems that define how frontier AI models are evaluated and secured.

You will iterate on our agent frameworks and build the scaffolding that lets models perform at their best, incorporating state-of-the-art tools for task delegation and complex multi-step tasks.

You will design and build evaluation environments, including challenges that measure a model's ability to evade detection.

You will work with expert red teamers, pentesters, and contracted security firms to ground these environments in real-world adversarial practice.

You will own the pipeline end to end, from designing APIs for frontier labs to building the analysis and visualization tools that summarize 10,000+ agent trajectories into specific conclusions.

From time to time, you will publish results in blog posts and academic venues, and deliver findings directly to customers.

You may be a good fit if
  • You have strong production programming skills and can move quickly from prototype to reliable infrastructure.
  • You have experience building agentic systems, evaluation harnesses, or environment frameworks, or are eager to learn.
  • You have security experience, or are excited to build it here.
  • You have strong problem-solving and analytical skills, and can extract clear signal from messy data.
You'll thrive here if
  • You are excited about applying novel research and new technology to real-world problems.
  • You are excited to learn new domains and build deep context in AI security.
  • You are results-oriented and can balance deep exploration with practical implementation.
  • You proactively step outside defined boundaries to support the team's mission and unblock others.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Post-Training
Research Engineer, Post-Training

General Analysis • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Forward Deployed Engineer
Forward Deployed Engineer

General Analysis • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
AI Security Research Engineer: Evaluation & Defense
AI Security Research Engineer: Evaluation & Defense

General Analysis • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Member of Technical Staff, AI Products
Member of Technical Staff, AI Products

Zaun • San Francisco (CA)

On-site
USD 150,000 - 210,000
AI security researcher
AI security researcher

0Labs • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Flexible work setup
Remote or hybrid work
+6
Research Engineer — AI Alignment & Evaluation
Research Engineer — AI Alignment & Evaluation

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Research Engineer, Frontier Evals & Environments
Research Engineer, Frontier Evals & Environments

OpenAI • Los Angeles (CA)

On-site
USD 100,000 - 150,000
Security Researcher
Security Researcher

QUORE IT : Talent Sourcing & Recruitment • New York (NY)

Hybrid
USD 250,000 - 300,000
Relocation support
Equity package
Research Engineer, Frontier Evals & Environments
Research Engineer, Frontier Evals & Environments

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Applied Research Scientist
Applied Research Scientist

Fleet AI, Inc. • Buffalo (NY)

On-site
USD 150,000 - 210,000