Research Engineer, Applied AI

HeyMilo AI

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments with verifiable rewards for real-world use cases, creating simulators, reward functions, and evaluation harnesses to measure model performance.

The role blends ML research with systems engineering, collaborating directly with founders and advisors to take a use case from problem definition to reproducible environments for training and evaluation.

Qualifications

  • Master's degree or PhD in AI, ML, CS or related field.
  • Solid grounding in reinforcement learning and LLM post-training (reward design, policy optimization, evaluation).
  • Strong software engineering skills in Python and reproducible code.

Responsibilities

  • Design and build RL environments for real-world use cases, including simulators and reward functions.
  • Create verifiable evaluation harnesses with clean scoring and cost tracking.
  • Run post-training experiments to validate learnable signals from environments.
  • Package environments and results for reproducibility and publish evaluations.
  • Contribute to research write-ups and determine next-use cases.

Skills

Reinforcement learning
LLM evaluation
Python
Experiment design
Research collaboration

Education

Master's degree or PhD in AI/ML/CS

Tools

Docker
Linux
Cloud environments

Job description

HeyMilo is building AI interviewers that automate and improve hiring through conversational AI. We work closely with companies to bring AI into real hiring workflows.

We also run an applied AI team that studies where today's models succeed and fail across industries.

The Role

We're hiring a Research Engineer to join our Applied AI team in San Francisco. You'll design and build reinforcement learning environments with verifiable rewards for specific real-world use cases: the simulators, reward functions, and evaluation harnesses that let us measure and improve how models perform on real work.

The work is equal parts ML research and systems engineering. You'll work directly with the founders and our research advisors, taking a use case from problem definition to a reproducible environment that models can be evaluated and trained against.

What You'll Do
  • Design and build RL environments for specific real-world use cases: realistic simulators, tool interfaces, and episodic task generation with proper isolation and reproducibility
  • Design verifiable reward functions that score correct intermediate actions as well as end states, and hold up against reward hacking
  • Build and maintain evaluation harnesses that run task suites across frontier and open-weight models, with clean scoring and cost tracking
  • Run post-training experiments (e.g. RLVR-style fine-tuning of open models) to validate that your environments produce a learnable signal
  • Package environments and results for reproducibility, and contribute to research write-ups and published evaluations
  • Help define which use cases we pursue next, informed by where models are weakest
What We're Looking For
  • Master's or PhD in AI, Machine Learning, Computer Science, or a closely related field
  • Solid grounding in reinforcement learning and LLM post-training (reward design, policy optimization, evaluation methodology)
  • Strong software engineering skills in Python, with code that others can run and build on
  • Hands‑on experience with LLMs: running evaluations, building agentic loops, tool calling, fine‑tuning
  • Comfortable with containers and infrastructure (Docker, Linux, cloud environments) for reproducible experiment setups
  • Ability to operate in ambiguity and move quickly; comfortable owning a problem end to end
  • Based in (or willing to relocate to) the San Francisco Bay Area
Bonus
  • Published research or open-source contributions in ML, RL, or evaluation
  • Experience with RL/eval frameworks and simulated or sandboxed environments
  • Experience training or fine-tuning open-weight models at any scale
Why join
  • Ground-floor role on a new applied AI team with real influence over how we build
  • Work directly with the founders and experienced research advisors, with your name on published work
  • High visibility, fast‑paced, execution‑driven environment
  • Competitive pay, equity, and benefits
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Research Engineer - RL Environments & Evaluation
Applied AI Research Engineer - RL Environments & Evaluation

HeyMilo AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer, Research Acceleration
Software Engineer, Research Acceleration

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Research Acceleration
Software Engineer, Research Acceleration

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer
Research Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 120,000 - 140,000
Health coverage
Opportunity to work with leading AI labs
Competitive salary and equity
Director, Applied Machine Learning
Director, Applied Machine Learning

Handshake • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Equity ownership
401(k) match
Parental leave
+3
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Director, Applied Machine Learning
Director, Applied Machine Learning

Apply • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity
401(k) match
Parental leave
+10
Research Engineer/Scientist - Human Alignment, Consumer Devices
Research Engineer/Scientist - Human Alignment, Consumer Devices

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 445,000
Research Scientist - Multimodal Agent, Consumer Devices
Research Scientist - Multimodal Agent, Consumer Devices

OpenAI • United States

Hybrid
USD 140,000 - 230,000
Relocation assistance
Hybrid work model (3 days in office)
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3