Benchmarking Software Engineer — Remote

Office Hours

United States

On-site

USD 160,000 - 210,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical/dental/vision
401(k)
Wellness stipend
Paid time off
Off-sites
Remote flexibility

Job summary

Office Hours is seeking a Software Engineer to build and run the platform behind our AI model evaluations. You will partner with researchers to prepare benchmark datasets, create repeatable evaluation pipelines, and deliver publishable results.

You will design containerized environments, support model experiments, and develop dashboards and leaderboards to track progress. This role offers remote or hybrid work from SF/NYC and a competitive package.

Qualifications

  • 4+ years of professional software engineering experience.
  • Strong Python and Docker experience.
  • Ability to work with researchers and engineers to translate methodology into systems.

Responsibilities

  • Prepare and maintain benchmark datasets and cleaning.
  • Build evaluation pipelines across model APIs and agents to ensure reproducibility.
  • Create containerized evaluation environments and viewers for tasks and model performance.
  • Support model experiments and compare baselines.
  • Develop the scoreboard and leaderboards for results across models and tasks.
  • Build analysis tools to identify failure modes and track improvements over time.
  • Collaborate with researchers and engineers to integrate data and outputs into publications.
  • Build lightweight HTML viewers and internal tools for task authoring and data collection.

Skills

Python
Docker
Data quality
Collaboration

Tools

Harbor
Terminal-Bench
Inspect

Job description

Office Hours is seeking a Software Engineer to build and run the platform behind our AI model evaluations. You will partner with researchers to prepare benchmark datasets, create repeatable evaluation pipelines, and deliver publishable results.

You will design containerized environments, support model experiments, and develop dashboards and leaderboards to track progress. This role offers remote or hybrid work from SF/NYC and a competitive package.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Benchmarking Software Engineer — Remote, Equity
Benchmarking Software Engineer — Remote, Equity

Office-Hours • San Francisco (CA), New York (NY)

Hybrid
USD 160,000 - 210,000
Competitive salary
Equity
Medical, dental, and vision coverage
+5
Software Engineer- Benchmarking
Software Engineer- Benchmarking

Office-Hours • San Francisco (CA), New York (NY)

Hybrid
USD 160,000 - 210,000
Competitive salary
Equity
Medical, dental, and vision coverage
+5
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
Remote Senior Software Engineer - AI-Driven Platform
Remote Senior Software Engineer - AI-Driven Platform

Benchmark Analytics • United States

On-site
USD 150,000 - 170,000
Unlimited Paid Time Off
Medical, dental, and vision benefits
401(k) plan
+1
Software Engineer, Benchmarking
Software Engineer, Benchmarking

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
AI-Driven Full-Stack Engineer - Remote, Ship Fast
AI-Driven Full-Stack Engineer - Remote, Ship Fast

Benchmark Analytics • United States

On-site
USD 120,000 - 140,000
Unlimited Paid Time Off
Fully remote environment
Medical, dental, and vision plans
+2
Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer — AI Benchmarking & Systems Infra
Software Engineer — AI Benchmarking & Systems Infra

LatchBio • San Francisco (CA)

On-site
USD 180,000 - 250,000
Unlimited PTO (truly)
Waterfront office in China Basin, San 
Free lunch and dinner
+1
AI Benchmarking Engineer — Evaluation & Failure Analysis
AI Benchmarking Engineer — Evaluation & Failure Analysis

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
End-to-End Platform Engineer for AI Benchmarking
End-to-End Platform Engineer for AI Benchmarking

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 210,000