Refresh AI is seeking a Research Engineer in San Francisco to push the boundaries of benchmarking technology. You will build benchmarks that labs use for evaluating coding abilities and computer-use capability. Your role will require expertise in reinforcement learning and supervised fine-tuning, as well as a willingness to engage in full-stack development on our core tech stack (Vercel, Supabase, Render). This position offers a full-time role in a dynamic environment.
Qualifications
Experience with reinforcement learning and supervised fine-tuning.
Ability to ship full-stack work on core technologies when required.
Responsibilities
Build benchmarks for coding and computer-use capability.
Translate expert workflows into rigorous evaluations.
Run evaluations against frontier models and publish verifiable numbers.
Skills
Reinforcement learning
Supervised fine-tuning
Full-stack development
Tools
Vercel
Supabase
Render
Job description
# Research Engineer, BenchmarkingEngineeringSan FranciscoFull-timeBuild the benchmarks frontier labs use to measure real-world coding and computer-use capability. Translate expert workflows into rigorous, verifiable evaluations, run them against frontier models, and publish numbers that hold up under adversarial scrutiny.Across every role at Refresh: you're willing to ship full-stack work on our core stack (Vercel, Supabase, Render) when it's needed, and you're comfortable with reinforcement learning and supervised fine-tuning at a high level.