We are looking for a Vision-Language-Action (VLA) Research Engineer to help build the next generation of intelligent robotic systems. In this role, you will work at the intersection of computer vision, language models, and embodied control, developing models that allow robots to perceive the world, reason with language, and take meaningful actions in real-world environments.
You will collaborate closely with systems engineers to translate cutting-edge research into scalable, real-world robotic capabilities.
Responsibilities
- Design, implement, and evaluate Vision-Language-Action models for embodied agents and robotic systems
- Develop multimodal learning pipelines combining visual perception, language understanding, and action/control
- Train and fine-tune large-scale models using simulation and real-world robotic data
- Explore approaches such as imitation learning, reinforcement learning, foundation models, and policy learning
- Integrate perception and decision-making models with robotic hardware and simulators
- Conduct experiments, analyze results, and iterate rapidly on model architectures
- Collaborate on research publications, internal reports, and technical documentation
- Stay current with the latest research in robotics, multimodal learning, and foundation models
Qualifications
- Strong background in machine learning, robotics, computer vision, or NLP
- Experience with deep learning frameworks (e.g., PyTorch, JAX, TensorFlow)
- Hands-on experience with vision-language models, policy learning, or embodied AI
- Solid programming skills in Python (C++ a plus)
- Familiarity with robotics concepts such as control, kinematics, sensors, or simulation
- Ability to read, implement, and extend research papers
Preferred Qualifications
- Experience training or deploying Vision-Language-Action or multimodal foundation models
- Experience with robotic simulators (e.g., Isaac Sim, MuJoCo, Habitat, Gazebo)
- Background in reinforcement learning, imitation learning, or offline RL
- Experience working with real robotic platforms (manipulation, navigation, mobile robots, etc.)
- Publications in top-tier conferences or journals (e.g., RSS, ICRA, CoRL, NeurIPS, ICML, CVPR)
- Experience scaling training pipelines on distributed or cloud systems