Multimodal Robotics Foundation Model Engineer

Lightwheel

California (MO)

On-site

USD 140,000 - 200,000

Full time

19 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lightwheel is seeking a research-focused engineer to advance vision-language-action models for real-world robotics. The role involves designing end-to-end training loops, coordinating with simulation and deployment teams, and driving robust policy deployment across robot platforms.

Ideal candidates hold a Master's or PhD with strong Python/PyTorch skills and experience in RL, multimodal models, and robot learning.

Qualifications

  • Master's or PhD in machine learning, robotics, computer vision, control, or a related field preferred.
  • Deep experience in at least one of the following areas: VLA models, robot learning, reinforcement learning, imitation learning, or multimodal foundation models.
  • Proficiency in Python and PyTorch; experience with CUDA/JAX, distributed training, or large-scale data pipelines is a plus.
  • Experience with real robots, simulation-based training, teleoperation/human data, or model deployment.
  • Able to clearly articulate personal contributions, training data, baselines, ablations, failure cases, and the limits of reported results.

Responsibilities

  • Develop vision-language-action (VLA) models and policies, Diffusion Policies, behavior cloning, offline/online reinforcement learning, and related training methods.
  • Design unified representations, data mixture strategies, and quality evaluation systems across human, robot, and simulation data.
  • Build an end-to-end training loop spanning pretraining, supervised fine-tuning (SFT), preference optimization, policy evaluation, and real-robot rollouts.
  • Research embodiment adaptation across robot platforms, grippers, and sensor configurations.
  • Collaborate with simulation, world model, data, and deployment teams to move models from experimentation into real-world robotic systems.
  • Establish quantitative metrics for model generalization, long-tail failures, and deployment stability.

Skills

Python
PyTorch
Reinforcement learning
Multimodal models
Robot learning

Education

Master's or PhD in ML/Robotics/CS

Tools

CUDA
JAX
C++
distributed training
large-scale data pipelines

Job description

Lightwheel is seeking a research-focused engineer to advance vision-language-action models for real-world robotics. The role involves designing end-to-end training loops, coordinating with simulation and deployment teams, and driving robust policy deployment across robot platforms.

Ideal candidates hold a Master's or PhD with strong Python/PyTorch skills and experience in RL, multimodal models, and robot learning.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multimodal Foundation Models Engineer for Robotics
Multimodal Foundation Models Engineer for Robotics

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
Lead Multimodal ML for Robotic Systems (Edge AI)
Lead Multimodal ML for Robotic Systems (Edge AI)

Pantera Capital • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff ML Engineer: Multimodal Perception & 3D World Models
Staff ML Engineer: Multimodal Perception & 3D World Models

Waymo • Mountain View (CA)

Hybrid
USD 251,000 - 310,000
Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

Pantera Capital • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
ML Engineer, Foundation Models — Embodied AI Recipes
ML Engineer, Foundation Models — Embodied AI Recipes

Waymo • San Francisco (CA)

Hybrid
USD 175,000 - 215,000
Annual bonus
Equity incentives
Company benefits
Robotics Foundation Model Engineer
Robotics Foundation Model Engineer

Lightwheel • California (MO)

On-site
USD 140,000 - 200,000
Staff Research Scientist, Multimodal Perception WorldModels
Staff Research Scientist, Multimodal Perception WorldModels

Neura Market • Mountain View (CA)

Hybrid
USD 251,000 - 310,000
Health insurance
401(k) with company match
Paid time off - 20 days/year
+4
Senior ML Engineer — 3D Perception & Multimodal AI
Senior ML Engineer — 3D Perception & Multimodal AI

Socket.dev • Mountain View (CA)

Hybrid
USD 213,000 - 263,000
Health insurance
401(k) with company match
Paid time off
ML Engineer for Multimodal Robotics & Embodied AI
ML Engineer for Multimodal Robotics & Embodied AI

Human Archive • San Francisco (CA)

On-site
USD 120,000 - 160,000