Benchmarking Research Engineer: Frontier Model Evaluations
Refresh AI
San Francisco (CA)
On-site
USD 120,000 - 150,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Refresh AI is seeking a Research Engineer in San Francisco to push the boundaries of benchmarking technology. You will build benchmarks that labs use for evaluating coding abilities and computer-use capability. Your role will require expertise in reinforcement learning and supervised fine-tuning, as well as a willingness to engage in full-stack development on our core tech stack (Vercel, Supabase, Render). This position offers a full-time role in a dynamic environment.
Qualifications
Experience with reinforcement learning and supervised fine-tuning.
Ability to ship full-stack work on core technologies when required.
Responsibilities
Build benchmarks for coding and computer-use capability.
Translate expert workflows into rigorous evaluations.
Run evaluations against frontier models and publish verifiable numbers.
Skills
Reinforcement learning
Supervised fine-tuning
Full-stack development
Tools
Vercel
Supabase
Render
Job description
Refresh AI is seeking a Research Engineer in San Francisco to push the boundaries of benchmarking technology. You will build benchmarks that labs use for evaluating coding abilities and computer-use capability. Your role will require expertise in reinforcement learning and supervised fine-tuning, as well as a willingness to engage in full-stack development on our core tech stack (Vercel, Supabase, Render). This position offers a full-time role in a dynamic environment.