Research Engineer, Benchmarking

Refresh AI

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Refresh AI is seeking a Research Engineer in San Francisco to push the boundaries of benchmarking technology. You will build benchmarks that labs use for evaluating coding abilities and computer-use capability. Your role will require expertise in reinforcement learning and supervised fine-tuning, as well as a willingness to engage in full-stack development on our core tech stack (Vercel, Supabase, Render). This position offers a full-time role in a dynamic environment.

Qualifications

  • Experience with reinforcement learning and supervised fine-tuning.
  • Ability to ship full-stack work on core technologies when required.

Responsibilities

  • Build benchmarks for coding and computer-use capability.
  • Translate expert workflows into rigorous evaluations.
  • Run evaluations against frontier models and publish verifiable numbers.

Skills

Reinforcement learning
Supervised fine-tuning
Full-stack development

Tools

Vercel
Supabase
Render

Job description

# Research Engineer, BenchmarkingEngineeringSan FranciscoFull-timeBuild the benchmarks frontier labs use to measure real-world coding and computer-use capability. Translate expert workflows into rigorous, verifiable evaluations, run them against frontier models, and publish numbers that hold up under adversarial scrutiny.Across every role at Refresh: you're willing to ship full-stack work on our core stack (Vercel, Supabase, Render) when it's needed, and you're comfortable with reinforcement learning and supervised fine-tuning at a high level.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Benchmarking Engineer — Evaluation & Failure Analysis
AI Benchmarking Engineer — Evaluation & Failure Analysis

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
AI Benchmarking & Evaluation Engineer
AI Benchmarking & Evaluation Engineer

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
Infra Engineer for Frontier AI Biology Benchmarks
Infra Engineer for Frontier AI Biology Benchmarks

LatchBio • San Francisco (CA)

On-site
USD 180,000 - 250,000
Unlimited PTO
Premium health plan
Office in San Francisco
+3
Senior Technical Researcher & Benchmark Writer — Remote, Equity
Senior Technical Researcher & Benchmark Writer — Remote, Equity

SearchApi • United States

Remote
USD 80,000 - 120,000
Fully Remote
Equity share
Profit sharing
+2
Remote AI Benchmark Engineer & Researcher
Remote AI Benchmark Engineer & Researcher

Pathway • Palo Alto (CA)

Remote
USD 120,000 - 180,000
Performance Platform Engineer — Equity & Benchmarking Infra
Performance Platform Engineer — Equity & Benchmarking Infra

Icehouseventures • Mountain View (CA)

On-site
USD 152,000 - 228,000
Software Engineer, Benchmarking
Software Engineer, Benchmarking

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
Research Engineer - Frontier AI Training & Evals
Research Engineer - Frontier AI Training & Evals

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Competitive compensation
Medical, dental, vision coverage
Lunch and dinner in office
+5
Staff Engineer, Paper Decomposition & Benchmarking
Staff Engineer, Paper Decomposition & Benchmarking

Infinity Artificial Intelligence Institute • San Francisco (CA)

On-site
USD 120,000 - 190,000