Principal AI Research Engineer - RL

reflexrobotics

New York (NY)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Reflex Robotics in New York designs AI-powered humanoid robots for tough industrial tasks. We are a small, mission-driven startup backed by top investors and real hardware with profitable unit economics.

We’re seeking an on-policy RL engineer to build robust robot policies, debug gradients, and push performance toward near-perfect deployment. You’ll own meaningful contributions and impact from day one, shaping the product’s future across real-world use cases.

Qualifications

  • You re-implemented core RL algorithms (SAC, DDPG) and can debug unstable gradients.
  • You contributed to sample-efficient RL methods (DreamerV3, MuZero).
  • You shipped on-policy RL on real hardware (e.g., robotics deployments).

Responsibilities

  • Develop robust on-policy RL policies for robotic systems.
  • Debug gradients and tune hyperparameters for stable learning.
  • Ship RL solutions on hardware in real-world environments.

Skills

RL reimplementation (SAC, DDPG)
DreamerV3 / MuZero contributions
On-policy RL deployment on hardware

Job description

Company Overview

Reflex Robotics builds general-purpose, AI-powered humanoid robots for the toughest jobs in industry.

We’re focused on jobs that are genuinely hard on the human body: lifting 40-pound bags of dog food, loading boxes for an entire shift, and moving heavy materials from floor to high overhead shelves. Not just lightweight tasks in controlled environments. Actual, back-breaking work.

Factory and warehouse work demands more. The hardware needs strength, reach, precision, speed, and the battery life to be useful across multiple shifts. It also needs to deliver a clear return today, not in some hypothetical future. Our robot costs $32K at development volumes today.

We build the full robot stack ourselves, from cameras and compute to actuators, electronics, and software. Our custom-built supervision stack helps handle the hard parts of deployment, from real floor variability to keeping work moving without line disruption. That supervision also generates proprietary real-world data that helps improve autonomy over time.

We are a four year old startup backed by Khosla Ventures and a team of other great investors. Our team includes alumni from Anduril, Oculus, Amazon, Waymo, Skydio, Boston Dynamics, ASML, MIT, Jane Street, and Joby Aviation.

Key Company Beliefs

  • We are obsessed with shipping robots.

  • A humanoid robot is only valuable if it solves a real customer problem. Cool demos are fun, but reliable work is what matters.

  • Low cost is what makes widespread deployment possible.

  • High-quality, proprietary robotics data is the foundation of the next generation of physical AI.

  • Getting nerd-sniped by an engineering metric is way less important than solving our customers’ biggest pain points.

What We’re Looking For

We’re looking for stellar on-policy RL engineers to work on creating robust robot policies.

We’re still a small team—which means high ownership, high equity, and the chance to shape the product from the ground up.

VLAs and other great “base policies” for robotics achieve ~80% success rates, but in real robot deployments, it’s essential to achieve 99.99% success rates. We can’t ask our customers to tolerate our robots packing three socks into a bin instead of four, or swapping shipping labels between two packages—not even once!
You should apply for this role if:

  • You’ve re-implemented core RL algorithms (SAC, DDPG) from scratch and can debug unstable gradients / tune hyperparameters correctly

  • You’ve made meaningful intellectual contributions to sample-efficient RL algorithms (e.g., DreamerV3 and MuZero)

  • You’ve shipped on-policy RL on hardware that learns in the real-world (e.g., for quadruped walking or drone racing)

You’d be joining a company that already has a solid core business—with working hardware, delighted customers, and profitable unit economics. Reflex is de-risked enough to see the hazy outlines of success, but still small enough that there’s enormous upside up for grabs.

Come Join Us

This is a rare opportunity to help build a flagship robotics company from the ground up—and to do work that will truly matter, reshaping what people believe is possible in robotics.

We love to see the things you’ve worked on. Have a portfolio or insane project you’ve worked on? Share it. We’re looking for people who push past the status quo, are passionate at work and in their own time—we’re looking for people who want to win.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Research Engineer - World Models
Principal AI Research Engineer - World Models

reflexrobotics • New York (NY)

On-site
USD 140,000 - 190,000
Senior On-Policy RL Engineer — Robotics Autonomy
Senior On-Policy RL Engineer — Robotics Autonomy

reflexrobotics • New York (NY)

On-site
USD 140,000 - 210,000
Senior Firmware Engineer
Senior Firmware Engineer

Reflex Robotics • New York (NY)

On-site
USD 200,000 - 250,000
Senior Firmware Engineer
Senior Firmware Engineer

reflexrobotics • New York (NY)

On-site
USD 200,000 - 250,000
AI Engineer (RL & WBC)
AI Engineer (RL & WBC)

Foundation Robotics Labs Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Market standard benefits (health, vision, dental, 401k)
Robotics Engineer/Researcher - Robot Learning - Imitation Learning, Foundation Models, RL
Robotics Engineer/Researcher - Robot Learning - Imitation Learning, Foundation Models, RL

Proception Inc. • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Hardware Technician
Hardware Technician

Reflex Robotics • New York (NY)

On-site
USD 65,000 - 95,000
AI Researcher - Reinforcement Learning
AI Researcher - Reinforcement Learning

1X • San Carlos (CA)

On-site
USD 200,000 - 300,000
Comprehensive medical, dental, and vision coverage
Generous paid time off and parental leave
401(k) plan with matching contributions
+1
AI Researcher - Robot Learning
AI Researcher - Robot Learning

Halodi Robotics • San Carlos (CA)

On-site
USD 200,000 - 300,000
Medical, dental, and vision coverage
401(k) plan with company match
Commuter benefits
+1
Senior AI Software Engineer, Reinforcement Learning
Senior AI Software Engineer, Reinforcement Learning

Agility Robotics • Fremont (CA)

Hybrid
USD 187,000 - 292,000
401(k) Plan
Equity