Software Engineer- Benchmarking

Office Hours

United States

Hybrid

USD 160,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary and equity
Medical, dental, and vision coverage
401(k)
Monthly wellness and fitness stipend
Paid time off and company holidays
Annual company off-sites
Remote flexibility

Job summary

Office Hours is seeking a Software Engineer to build and run the platform behind our AI model evaluations. You’ll prepare benchmark datasets, construct reproducible evaluation pipelines, and create containerized environments to run models and tools.

Work closely with researchers to translate methodology into scalable systems. You will help fine-tune small models, develop leaderboards, and craft analysis tools to identify failure modes and track improvements over time, with a flexible remote or

Qualifications

  • 4+ years of professional experience building and maintaining complex systems.
  • Strong Python programming skills and code quality discipline.
  • Experience preparing, cleaning, and maintaining benchmark datasets to ensure comparability.
  • Ability to work with researchers to translate methodology into production-ready systems.

Responsibilities

  • Prepare and maintain benchmark datasets for evaluations.
  • Build and maintain evaluation pipelines for consistent results.
  • Create containerized evaluation environments and viewers for tasks.
  • Support fine-tuning and comparison of models and baselines.
  • Develop the scoreboard and leaderboard views for multi-level results.
  • Build tools to identify failure modes and track improvements over time.
  • Collaborate with researchers and engineers to ensure accurate, integrated outputs.
  • Build lightweight tooling for data creation and review.

Skills

Python
Data rigor
Docker
Collaborative

Tools

Docker

Job description

Software Engineer, Benchmarking (SF, NYC, or Remote)
About Us

Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. Experts earn income by sharing their knowledge through advisory work, projects, and AI model training. Our platform handles the complexities behind the scenes— screening, compliance, scheduling, and payments—so knowledge sharing stays focused on meaningful insights and real impact.

We’re a hyper-growth and profitable company, quickly expanding our expert network, launching new offices, and new products. We are headquartered in San Francisco, with offices in Brooklyn and Bangalore. Our customers include the fastest-growing digital health companies, technology companies, institutional investment firms, consulting firms and AI Labs. We are backed by top marketplace investors and operators of companies like DoorDash, Airbnb, affirm.

What we believe

Human knowledge is the world’s most valuable asset. And yet, despite being more interconnected than ever, most knowledge still remains stuck in our heads, inaccessible and underutilized. Our vision is to make human knowledge easily accessible and infinitely scalable by building tools for the new age knowledge economy.

About the role

We’re looking for a Software Engineer to build and run the platform behind our AI model evaluations. You’ll work closely with our research team to prepare benchmark datasets, build the pipelines and environments our evaluations run in, and turn results into published output.

Our researchers design the methodology. You’ll turn it into systems that run consistently and reproducibly, so results stay comparable across models, agent scaffolds, and time. Evaluations draw on the knowledge domains our expert network covers.

What you’ll do
  • Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results.

  • Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable.

  • Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use.

  • Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance.

  • Develop the scoreboard and leaderboard: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance.

  • Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time.

  • Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish.

  • Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.

What you bring
  • Solid engineering skills: 4+ years of professional experience building and maintaining complex systems, with strong Python. You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure.

  • Data rigor: Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable.

  • Comfort with containers and environments: Experience with Docker and building reproducible execution environments.

  • Collaborative: You work well alongside researchers and scientists and can translate their methodology into working systems.

Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus.

Tech Stack
  • Evaluations: Python, model APIs, agent/evaluation frameworks, custom evaluation tooling

  • Models: APIs from the major AI providers, terminal agents, and open-source models via the Hugging Face ecosystem and PyTorch

  • Environments: Docker

  • Publishing: React, Next.js, Tailwind

  • Workflow: GitHub, Slack, Notion, Linear

Bonus Experience
  • Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work

  • Experience with agentic, multi-turn, long-context, or tool-use evaluation

  • Experience validating LLM-as-judge or rubric-based grading setups

  • Background or strong interest in a scientific or technical domain

  • Experience building data-heavy dashboards, leaderboards, or visualizations

  • Open-source contributions or published work related to benchmarks and measurement

Benefits + Perks
  • Competitive salary and equity

  • Medical, dental, and vision coverage

  • 401(k)

  • Monthly wellness and fitness stipend

  • Paid time off policy, along with company holidays

  • Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City)

  • Parent-friendly policies, remote flexibility, and paid family leave

Pay Transparency Notice

Full-time offers include base salary, equity, and benefits.

Pay range: $160,000-$210,000, based on seniority, relevant experience and location

This role can be fully remote or hybrid out of our SF or NYC offices.

Don’t meet every single requirement? Studies have shown that some candidates, especially underrepresented groups such as women and people of color, are less likely to apply to jobs unless they meet every single qualification. At Office Hours we believe in building a diverse and inclusive workplace, so if you’re excited about this role but don’t meet every qualification in the job description, we still encourage you to apply. You could still be the right candidate for this or other roles at Office Hours!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer- Benchmarking
Software Engineer- Benchmarking

Office-Hours • San Francisco (CA), New York (NY)

Hybrid
USD 160,000 - 210,000
Competitive salary
Equity
Medical, dental, and vision coverage
+5
Engagement Manager (Expert Data Services)
Engagement Manager (Expert Data Services)

Office Hours • San Francisco (CA)

Hybrid
USD 130,000 - 180,000
Competitive salary and equity
Medical, dental, and vision coverage
401(k)
+4
Strategic Account Executive, AI Labs
Strategic Account Executive, AI Labs

Office Hours • United States

Hybrid
USD 275,000 - 350,000
Medical coverage
Dental coverage
Vision coverage
+6
Software Engineer, Full Stack
Software Engineer, Full Stack

Office Hours • San Francisco (CA)

On-site
USD 165,000 - 185,000
Medical, dental, and vision coverage
401(k)
Wellness stipend
+2
Software Engineer, Full Stack
Software Engineer, Full Stack

Office Hours • New York (NY)

On-site
USD 165,000 - 185,000
Competitive salary and equity
Medical, dental, and vision coverage
401(k)
+4
Software Engineer, Full Stack
Software Engineer, Full Stack

Office Hours • United States

On-site
USD 165,000 - 185,000
Medical, dental, and vision coverage
401(k)
Wellness and fitness stipend
+3
Recruiter, San Francisco
Recruiter, San Francisco

Office Hours • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive salary & stock options
Healthcare, dental, and vision coverage
Wellness/fitness benefit
+6
People Operations Lead
People Operations Lead

Socket.dev • New York (NY), San Francisco (CA)

Hybrid
USD 140,000 - 160,000
Competitive salary & equity
Medical, dental, vision coverage
401(k)
+3
Client Solutions Manager
Client Solutions Manager

SupportFinity™ • New York (NY)

Hybrid
USD 110,000 - 130,000
Stock options
Healthcare
Dental & Vision
+5
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Socket.dev • Beverly Hills (CA)

On-site
USD 180,000 - 280,000
Daily team dinner provided in-office