Research Engineer

Confero

San Francisco (CA)

Hybrid

USD 200,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Confero in San Francisco is seeking a Research Engineer, AI Safety / Alignment to join a hybrid team working three days in the office. You will design and build evaluation environments and RL-style tests to study how frontier models behave when optimized for the wrong objective.

You will own projects end to end, from scenarios to agents and graders, and contribute to infrastructure that runs and evaluates these experiments.

Qualifications

  • Strong generalist software engineering ability with Python preference.
  • At least 1 year of professional software engineering experience.
  • Genuine interest in AI alignment and reducing existential risk.
  • Strong analytical and conceptual reasoning.
  • Comfortable owning ambiguous technical projects end to end.

Responsibilities

  • Design and build evaluation and RL-style environments to test frontier models for reward hacking.
  • Own research projects end to end: developing scenarios, building environments, directing agents, iterating graders, analyzing behavior.
  • Contribute to shared infrastructure used to create, run and evaluate these environments.

Skills

Python
Software engineering
Analytical thinking
Project ownership
AI alignment interest

Tools

RL frameworks

Job description

Research Engineer, AI Safety / Alignment
San Francisco | Hybrid, 3 days in office
$200K-$300K base + equity

We’re building a new applied AI research team focused on AI alignment and existential risk, studying how increasingly capable models and agents behave when given opportunities to exploit, cheat or optimise for the wrong objective.

What you’ll do

You’ll design and build evaluation and RL-style environments that test frontier models for behaviours such as reward hacking and exploiting weaknesses in tasks or graders.

You’ll own research projects end to end: developing scenarios, building environments, directing agents, iterating graders, analysing model behaviour and refining the experiment.

You’ll also contribute to the shared infrastructure used to create, run and evaluate these environments.

Prior experience with RL environments, agent evaluations or reinforcement learning is valuable, but not required.

A good fit:
  • Strong generalist software engineering ability, ideally Python
  • 1+ years of professional software engineering experience
  • Genuine interest in AI alignment and reducing existential risk from advanced AI
  • Strong analytical and conceptual reasoning
  • Comfortable owning ambiguous technical projects end to end
  • Interested in working directly with frontier LLMs and AI agents
  • Experience with evals, agents, RL environments, security or adversarial systems is a plus
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer — AI Alignment & Evaluation
Research Engineer — AI Alignment & Evaluation

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
AI Safety Research Engineer - RL Environments & Alignment
AI Safety Research Engineer - RL Environments & Alignment

Confero • San Francisco (CA)

Hybrid
USD 200,000 - 300,000
Research Engineer, AI Safety & Evaluations
Research Engineer, AI Safety & Evaluations

TTN Talent • San Francisco (CA)

On-site
USD 200,000 - 400,000
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Relocation assistance
Research Scientist, AI Alignment & Safety
Research Scientist, AI Alignment & Safety

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Research Scientist, Alignment
Research Scientist, Alignment

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 500,000
Relocation assistance
Hybrid work model
AI Alignment Research Engineer — Evaluation & Safety
AI Alignment Research Engineer — Evaluation & Safety

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Research Engineer – RL Infrastructure & Agent Environments
Research Engineer – RL Infrastructure & Agent Environments

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 190,000
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours