AI Safety Research Engineer - RL Environments & Alignment

Confero

San Francisco (CA)

Hybrid

USD 200,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Confero in San Francisco is seeking a Research Engineer, AI Safety / Alignment to join a hybrid team working three days in the office. You will design and build evaluation environments and RL-style tests to study how frontier models behave when optimized for the wrong objective.

You will own projects end to end, from scenarios to agents and graders, and contribute to infrastructure that runs and evaluates these experiments.

Qualifications

  • Strong generalist software engineering ability with Python preference.
  • At least 1 year of professional software engineering experience.
  • Genuine interest in AI alignment and reducing existential risk.
  • Strong analytical and conceptual reasoning.
  • Comfortable owning ambiguous technical projects end to end.

Responsibilities

  • Design and build evaluation and RL-style environments to test frontier models for reward hacking.
  • Own research projects end to end: developing scenarios, building environments, directing agents, iterating graders, analyzing behavior.
  • Contribute to shared infrastructure used to create, run and evaluate these environments.

Skills

Python
Software engineering
Analytical thinking
Project ownership
AI alignment interest

Tools

RL frameworks

Job description

Confero in San Francisco is seeking a Research Engineer, AI Safety / Alignment to join a hybrid team working three days in the office. You will design and build evaluation environments and RL-style tests to study how frontier models behave when optimized for the wrong objective.

You will own projects end to end, from scenarios to agents and graders, and contribute to infrastructure that runs and evaluates these experiments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Confero • San Francisco (CA)

Hybrid
USD 200,000 - 300,000
AI Alignment Research Engineer — Evaluation & Safety
AI Alignment Research Engineer — Evaluation & Safety

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Senior ML Engineer: AI Safety & Alignment (RLHF)
Senior ML Engineer: AI Safety & Alignment (RLHF)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
AI Alignment Research Scientist — Safety & Reasoning
AI Alignment Research Scientist — Safety & Reasoning

Safetytalent • San Francisco (CA)

On-site
USD 100,000 - 150,000
Research Scientist, AI Alignment & Safety
Research Scientist, AI Alignment & Safety

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Research Engineer — AI Alignment & Evaluation
Research Engineer — AI Alignment & Evaluation

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Frontier-Model Safety Researcher & Evaluations
Frontier-Model Safety Researcher & Evaluations

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 500,000
Relocation assistance
Hybrid work model
Frontier AI Safety Researcher (Hybrid — SF)
Frontier AI Safety Researcher (Hybrid — SF)

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Relocation assistance
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Engineer, AI Safety & Evaluations
Research Engineer, AI Safety & Evaluations

TTN Talent • San Francisco (CA)

On-site
USD 200,000 - 400,000