About HUD
HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We've raised $16M from top VCs and were YC W25.
About the role
We're looking for a Full-Stack Software Engineer, Reinforcement Learning to build the product surfaces, backend systems, and internal tools that power HUD's RL data engine.
Responsibilities
- Develop product-facing tools for browsing environments, inspecting trajectories, reviewing task quality, debugging failures, and understanding model behavior
- Build vendor-facing workflows that make it easy for external partners to create, submit, test, and iterate on RL environments and training data
- Create dashboards and observability tools that surface environment quality, eval results, data collection progress, grader issues, reward signal problems, and pipeline health
- Design backend services and APIs that connect task authoring, data collection, evaluation, QA/QC, and RL training infrastructure
- Partner closely with research, operations, and GTM teams to turn vague, high-stakes requests into well-designed systems that ship quickly
Experience
You may be a good fit if you have:
- Strong software engineering fundamentals and real full-stack range, including proficiency in Python and a modern web stack such as React, TypeScript, Next.js, or similar
- Experience owning user-facing or internal products end-to-end
- Good product taste and the ability to build tools that are intuitive for both technical and non-technical users
- Comfort with cloud infrastructure, Docker, CI/CD, observability, and production debugging
- High agency-you identify what needs to exist, build it, and improve it without waiting for a perfect spec
- Strong communication skills for working across research, engineering, operations, vendors, and founders
Strong candidates may also have:
- Experience building data collection, labeling, annotation, eval, or research tooling platforms
- Experience building dashboards, review workflows, observability tools, or debugging interfaces for complex systems
- Experience building developer tools, infrastructure products, internal platforms, or workflow products that made a team dramatically faster
- Experience with AWS, Kubernetes, Terraform, Docker, Grafana, or similar infrastructure tools as tools to ship product, not as the center of the role
Team & company details
Team Size : ~15 people currently, mostly full-time in-person, but some remote.
Our team: Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.
Company stage: We have 8 figures in funding and high revenue growth. We're scaling profitably and quickly to meet very strong demand.
Logistics
- Employment : Full-time.
- Location : On-site only, for now. You can join the team in the San Francisco Bay Area or Singapore offices.
- Visa Sponsorship : We provide support for relocation and visas for strong full-time candidates to the US or Singapore.
- Timeline : Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.
What we offer
- Competitive compensation
- 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
- Lunch and dinner when you're in the office
- Company-wide holiday break (Christmas Eve to New Year's Day) on top of PTO and paid holidays
- Other perks including an Equinox membership, 401k, and commuter benefits (US employees)
- Unlimited* access to tokens for ChatGPT, Claude Code, Cursor, etc. * By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.