A complete application in a minute — tailored resume and cover letter, ready to send.
Confero in San Francisco is seeking a Research Engineer, AI Safety / Alignment to join a hybrid team working three days in the office. You will design and build evaluation environments and RL-style tests to study how frontier models behave when optimized for the wrong objective.
You will own projects end to end, from scenarios to agents and graders, and contribute to infrastructure that runs and evaluates these experiments.
We’re building a new applied AI research team focused on AI alignment and existential risk, studying how increasingly capable models and agents behave when given opportunities to exploit, cheat or optimise for the wrong objective.
You’ll design and build evaluation and RL-style environments that test frontier models for behaviours such as reward hacking and exploiting weaknesses in tasks or graders.
You’ll own research projects end to end: developing scenarios, building environments, directing agents, iterating graders, analysing model behaviour and refining the experiment.
You’ll also contribute to the shared infrastructure used to create, run and evaluate these environments.
Prior experience with RL environments, agent evaluations or reinforcement learning is valuable, but not required.