Research Engineer (General)

Visa Hunt

San Francisco, Northern (CA, KY)

On-site

USD 140,000 - 200,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Medical, dental, vision coverage
Lunch and dinner in office
Holiday break
Equinox membership
401k
Commuter benefits
Unlimited tokens for AI tools

Job summary

HUD is seeking Research Engineers to build the core systems for training and evaluating frontier AI agents. You will develop environments, improve data quality, and translate real workflows into benchmarks and tasks for RL training.

The role emphasizes constructing end-to-end pipelines, experiments, and tools for researchers and data vendors. On-site in the San Francisco Bay Area, with visa sponsorship for strong candidates.

Qualifications

  • Proficiency in Python, Docker, and Linux environments.
  • Experience with benchmarks and evaluation tasks for RL training.
  • Strong attention to detail and ability to spot inconsistencies in data or task design.
  • Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap.

Responsibilities

  • Build systems for creating, running, evaluating, and improving agent training environments.
  • Design experiments to understand model behavior, failure modes, and data quality issues.
  • Develop tools to help researchers and engineers create higher-quality tasks and trajectories.
  • Support the full lifecycle of agent training data from task design to validation.
  • Collaborate with external vendors to improve data engine throughput and quality.
  • Create metrics and analyses to assess usefulness of tasks and environments.

Skills

Python
Benchmarks & evals
Attention to detail
Remote collaboration

Tools

Docker
Linux

Job description

About HUD

HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.

About the role

This is a general application for candidates who are unsure which research focus - QC Automation, Benchmarks, or Synthetic Data - they would be a fit for. We would love to meet you and figure it out together. However, if you already have a focus in mind, please apply to only that application.

We\'re looking for Research Engineers to build the technical foundation for training and evaluating frontier AI agents. You’ll build the systems for creating new environments, improve data quality, and translate real-world workflows into tasks and benchmarks.

Responsibilities
  • Build systems for creating, running, evaluating, and improving agent training environments

  • Design experiments to understand model behavior, agent failure modes, and data quality issues

  • Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops

  • Work across the full lifecycle of agent training data - task design, environment setup, trajectory collection, evaluation, and validation

  • Partner with external vendors to identify bottlenecks and improve the quality and throughput of HUD’s data engine

  • Build metrics and analyses that help us understand whether our tasks, environments, and evals are actually useful for training frontier agents

Experience

You may be a good fit if you have:

  • Proficiency in Python, Docker, and Linux environments

  • Experience working on benchmarks and evals - you can reason about what makes a task realistic, a rubric reliable, an environment usable, and a trajectory useful for RL training

  • Strong attention to detail and the ability to spot subtle inconsistencies in data, model behavior, or task design

  • Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap

  • Early-stage startup experience with ability to work independently in fast-paced environments

Strong candidates may also have:

  • Experience building internal tools, research infrastructure, or data pipelines

  • Experience designing metrics and validation workflows

  • A background in competitive programming, Olympiad medaling, research, or unusually strong independent project experience

  • Thrive in unstructured problem spaces

  • Strong communication skills for remote collaboration across time zones

We prioritize technical aptitude and learning potential over years of experience. Motivated candidates are encouraged to apply even if they don\'t meet all criteria.

Team & company details
  • Team Size: ~15 people currently, mostly full-time in-person, but some remote.

  • Our team: Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.

  • Company stage: We have 8 figures in funding and high revenue growth. We’re scaling profitably and quickly to meet very strong demand.

Logistics
  • Employment: Full-time.

  • Location: On-site in the San Francisco Bay Area.

  • Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US.

  • Timeline: Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.

What we offer
  • Competitive compensation based on experience and location

  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA

  • Lunch and dinner when you\'re in the office

  • Company-wide holiday break (Christmas Eve to New Year\’s Day) on top of PTO and paid holidays

  • Other perks including an Equinox membership, 401k, and commuter benefits

  • Unlimited* access to tokens for ChatGPT, Claude Code, Cursor, etc. *By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer (General)
Research Engineer (General)

HUD • San Francisco (CA)

On-site
USD 120,000 - 180,000
Blue Shield medical, dental, vision
Lunch and dinner in office
Holiday break
+4
Lead Research Engineer, Data Quality
Lead Research Engineer, Data Quality

PVH (Tommy Hilfiger/Calvin Klein) • San Francisco (CA)

Hybrid
USD 190,000 - 280,000
Medical, dental, vision (100% covered)
Lunch and dinner in the office
Equinox membership
+3
Research Engineer, QC Automation
Research Engineer, QC Automation

HUD • San Francisco (CA)

On-site
USD 140,000 - 200,000
Medical, dental, vision coverage (Blue
Lunch & dinner in office
Holiday break & PTO/holidays
+4
Marketing & Events Lead
Marketing & Events Lead

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Medical coverage
Dental coverage
Vision coverage
+6
Recruiter
Recruiter

hud • San Francisco (CA)

On-site
USD 80,000 - 120,000
100% covered medical, dental, and vision
Lunch and dinner in the office
Company-wide holiday break
+4
Marketing & Events Lead
Marketing & Events Lead

HUD • San Francisco (CA)

On-site
USD 80,000 - 120,000
100% covered medical, dental, and vision
Lunch and dinner in the office
Company-wide holiday break
+4
Growth Lead
Growth Lead

HUD • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
Top-tier medical/dental/vision
Lunch and dinner in office
+5
GTM Engineer
GTM Engineer

HUD • San Francisco (CA)

On-site
USD 120,000 - 180,000
Top-tier health insurance
Lunch & dinner in office
Holiday break
+4
Research Engineer
Research Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health coverage
Opportunity to work with leading AI labs
Competitive salary and equity
Software Engineer, Research Acceleration
Software Engineer, Research Acceleration

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1