AI Robotics Engineer, Vision-Language-Action (VLA)

Confidential

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

10 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Confidential is seeking a Vision-Language-Action (VLA) Research Engineer to advance intelligent robotic systems by integrating computer vision, language models, and embodied control. The role emphasizes collaboration with systems engineers to translate research into deployable robotic capabilities and scalable infrastructure.

The candidate will design multimodal models, train large-scale systems, and publish results while staying current with robotics and multimodal learning research.

Qualifications

  • Strong background in machine learning, robotics, computer vision, or NLP.
  • Hands-on experience with PyTorch, JAX, TensorFlow.
  • Solid programming skills in Python (C++ a plus).
  • Familiarity with robotics concepts such as control, kinematics, sensors, or simulation.
  • Ability to read, implement, and extend research papers.

Responsibilities

  • Design, implement, and evaluate Vision-Language-Action models for embodied agents and robotic systems.
  • Develop multimodal learning pipelines combining visual perception, language understanding, and action/control.
  • Train and fine-tune large-scale models using simulation and real-world robotic data.
  • Explore imitation learning, reinforcement learning, foundation models, and policy learning.

Skills

Machine learning
Robotics
Computer vision
NLP
Python

Tools

PyTorch
JAX
TensorFlow

Job description

We are looking for a Vision-Language-Action (VLA) Research Engineer to help build the next generation of intelligent robotic systems. In this role, you will work at the intersection of computer vision, language models, and embodied control, developing models that allow robots to perceive the world, reason with language, and take meaningful actions in real-world environments.

You will collaborate closely with systems engineers to translate cutting-edge research into scalable, real-world robotic capabilities.

Responsibilities
  • Design, implement, and evaluate Vision-Language-Action models for embodied agents and robotic systems
  • Develop multimodal learning pipelines combining visual perception, language understanding, and action/control
  • Train and fine-tune large-scale models using simulation and real-world robotic data
  • Explore approaches such as imitation learning, reinforcement learning, foundation models, and policy learning
  • Integrate perception and decision-making models with robotic hardware and simulators
  • Conduct experiments, analyze results, and iterate rapidly on model architectures
  • Collaborate on research publications, internal reports, and technical documentation
  • Stay current with the latest research in robotics, multimodal learning, and foundation models
Qualifications
  • Strong background in machine learning, robotics, computer vision, or NLP
  • Experience with deep learning frameworks (e.g., PyTorch, JAX, TensorFlow)
  • Hands-on experience with vision-language models, policy learning, or embodied AI
  • Solid programming skills in Python (C++ a plus)
  • Familiarity with robotics concepts such as control, kinematics, sensors, or simulation
  • Ability to read, implement, and extend research papers
Preferred Qualifications
  • Experience training or deploying Vision-Language-Action or multimodal foundation models
  • Experience with robotic simulators (e.g., Isaac Sim, MuJoCo, Habitat, Gazebo)
  • Background in reinforcement learning, imitation learning, or offline RL
  • Experience working with real robotic platforms (manipulation, navigation, mobile robots, etc.)
  • Publications in top-tier conferences or journals (e.g., RSS, ICRA, CoRL, NeurIPS, ICML, CVPR)
  • Experience scaling training pipelines on distributed or cloud systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vision-Language-Action Robotics Research Engineer
Vision-Language-Action Robotics Research Engineer

Confidential • San Francisco (CA)

On-site
USD 150,000 - 230,000
Robotics Engineer
Robotics Engineer

IFG - International Financial Group • Redmond (WA)

On-site
USD 140,000 - 180,000
Senior Research Scientist
Senior Research Scientist

Sereact • Massachusetts

On-site
USD 150,000 - 210,000
Health Insurance
401(k) with company match
20 days PTO
+5
Senior Research Scientist
Senior Research Scientist

Sereact GmbH • Boston (MA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Medical Insurance
Dental Insurance
Vision Insurance
+7
Member of Technical Staff, Vision/Language
Member of Technical Staff, Vision/Language

XDOF • San Mateo (CA)

On-site
USD 130,000 - 210,000
Senior Robotics Engineer - R&D
Senior Robotics Engineer - R&D

Sereact GmbH • Boston (MA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Medical/Dental/Vision
401k match
PTO 20 days
+5
Senior Robotics Engineer - R&D
Senior Robotics Engineer - R&D

Sereact • Boston (MA)

On-site
USD 150,000 - 225,000
Medical, dental, and vision insurance
401(k) with 100% company match, up to
20 days paid time off
+5
Robotics Software Engineer Vision-Language-Action Models
Robotics Software Engineer Vision-Language-Action Models

Trener Robotics • San Jose (CA)

On-site
USD 140,000 - 210,000
Research Engineer
Research Engineer

Acceler8 Talent • Mountain View (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Vision / Language
Member of Technical Staff, Vision / Language

xdof.ai • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and equity
Comprehensive health and wellness benefits
Collaborative and fast-paced work environment