Research Engineer (General)

HUD

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Blue Shield medical, dental, vision
Lunch and dinner in office
Holiday break
Equinox membership
401k
Commuter benefits
Unlimited access to AI tokens

Job summary

HUD is actively building infrastructure to create RL training data and evals for frontier AI agents, with a focus on a broad application. The role is a general application for candidates interested in QC Automation, Benchmarks, or Synthetic Data, with visa sponsorship for strong US candidates.

This is a full-time on-site position in the San Francisco Bay Area. You will build systems, design experiments, and develop tools across the data lifecycle, collaborating with researchers, engineers, and

Qualifications

  • Proficiency in Python and Linux environments.
  • Experience with benchmarks and evals; assess task realism and data quality.
  • Attention to data inconsistencies in models or tasks.
  • Experience building tools, pipelines, experiments, or infra.
  • Startup experience; able to work independently in fast-paced settings.

Responsibilities

  • Build systems for creating, running, evaluating, and improving agent training environments
  • Design experiments to understand model behavior, failure modes, and data quality
  • Develop tools to help researchers, engineers, and data vendors create better tasks, trajectories, and feedback loops
  • Work across the full lifecycle of agent training data from task design to evaluation
  • Partner with external vendors to improve data engine quality and throughput
  • Build metrics to assess usefulness of tasks, environments, and evals

Skills

Python
Linux
Attention to detail
Independent work

Tools

Docker
Data pipelines
Experimentation infra

Job description

About HUD

HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.

About The Role

This is a general application for candidates who are unsure which research focus - QC Automation, Benchmarks, or Synthetic Data - they would be a fit for. We would love to meet you and figure it out together. However, if you already have a focus in mind, please apply to only that application.

Responsibilities
  • Build systems for creating, running, evaluating, and improving agent training environments
  • Design experiments to understand model behavior, agent failure modes, and data quality issues
  • Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops
  • Work across the full lifecycle of agent training data - task design, environment setup, trajectory collection, evaluation, and validation
  • Partner with external vendors to identify bottlenecks and improve the quality and throughput of HUD’s data engine
  • Build metrics and analyses that help us understand whether our tasks, environments, and evals are actually useful for training frontier agents
Experience

You may be a good fit if you have:

  • Proficiency in Python, Docker, and Linux environments
  • Experience working on benchmarks and evals - you can reason about what makes a task realistic, a rubric reliable, an environment usable, and a trajectory useful for RL training
  • Strong attention to detail and the ability to spot subtle inconsistencies in data, model behavior, or task design
  • Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap
  • Early-stage startup experience with ability to work independently in fast-paced environments
Strong Candidates May Also Have
  • Experience building internal tools, research infrastructure, or data pipelines
  • Experience designing metrics and validation workflows
  • A background in competitive programming, Olympiad medaling, research, or unusually strong independent project experience
  • Thrive in unstructured problem spaces
  • Strong communication skills for remote collaboration across time zones
Team & Company Details
  • Team Size: ~15 people currently, mostly full-time in-person, but some remote.
  • Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.
  • Company stage: We have 8 figures in funding and high revenue growth. We’re scaling profitably and quickly to meet very strong demand.
Logistics
  • Employment: Full-time.
  • Location: On-site in the San Francisco Bay Area.
  • Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US.
  • Timeline: Applications are rolling. The process is 2 technical interviews and a 2‑3 day work trial.
What we offer
  • Competitive compensation based on experience and location
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA
  • Lunch and dinner when you’re in the office
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Other perks including an Equinox membership, 401k, and commuter benefits
  • Unlimited access to tokens for ChatGPT, Claude Code, Cursor, etc. By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.

Due to high volume, we may not actively respond to every application, but feel free to contact us at email address or elsewhere if we missed your application!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, QC Automation
Research Engineer, QC Automation

HUD • San Francisco (CA)

On-site
USD 140,000 - 200,000
Medical, dental, vision coverage (Blue
Lunch & dinner in office
Holiday break & PTO/holidays
+4
Research Engineer
Research Engineer

HUD • San Francisco (CA)

On-site
USD 140,000 - 200,000
Competitive compensation
Medical/Dental/Vision coverage
Lunch & dinner in office
+5
Research Engineer (General)
Research Engineer (General)

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Competitive compensation
Medical, dental, vision coverage
Lunch and dinner in office
+5
Lead Research Engineer, Data Quality
Lead Research Engineer, Data Quality

PVH (Tommy Hilfiger/Calvin Klein) • San Francisco (CA)

Hybrid
USD 190,000 - 280,000
Medical, dental, vision (100% covered)
Lunch and dinner in the office
Equinox membership
+3
Growth Lead
Growth Lead

HUD • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
Top-tier medical/dental/vision
Lunch and dinner in office
+5
Recruiter
Recruiter

hud • San Francisco (CA)

On-site
USD 80,000 - 120,000
100% covered medical, dental, and vision
Lunch and dinner in the office
Company-wide holiday break
+4
Marketing & Events Lead
Marketing & Events Lead

HUD • San Francisco (CA)

On-site
USD 80,000 - 120,000
100% covered medical, dental, and vision
Lunch and dinner in the office
Company-wide holiday break
+4
Marketing & Events Lead
Marketing & Events Lead

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Medical coverage
Dental coverage
Vision coverage
+6
GTM Engineer
GTM Engineer

HUD • San Francisco (CA)

On-site
USD 120,000 - 180,000
Top-tier health insurance
Lunch & dinner in office
Holiday break
+4
Research Engineer - Frontier AI Training & Evals
Research Engineer - Frontier AI Training & Evals

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Competitive compensation
Medical, dental, vision coverage
Lunch and dinner in office
+5