Staff ML Engineer, VLA Training

Nidus

Northern, New York (KY, NY)

Hybrid

USD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nidus Technologies is seeking a Staff Machine Learning Engineer to lead development of vision-language-action models powering our autonomous robots. You will own dataset construction, training recipes, and scalable training infrastructure, collaborating with researchers, robotics engineers, and data teams to deploy robust robot capabilities.

This hands-on role sits at the intersection of ML research, large-scale systems, and robotics, requiring strong communication and a track record of

Qualifications

  • Five+ years of professional experience in ML, robotics, computer vision, or related field.
  • Strong experience training modern deep learning models using PyTorch/JAX or similar frameworks.
  • Experience with transformer-based and multimodal models, or learned control policies.
  • Experience building and operating distributed training pipelines on multi-GPU/ multi-node clusters.
  • Experience with large-scale datasets, high-throughput data loading, experiment tracking, and reproducible ML workflows.
  • Strong understanding of optimization, model architecture, data quality, evaluation methodology, and experimental design.
  • Ability to debug failures across the full stack from data to robot behavior.
  • Bachelor's degree in CS/ML/Robotics or equivalent practical experience.

Responsibilities

  • Train and improve vision-language-action models for robotic manipulation and embodied tasks.
  • Develop model architectures and training recipes for multimodal policies.
  • Build scalable pipelines for pretraining, fine-tuning, evaluation, and model release.
  • Train models across multi-node GPU clusters, improving utilization and cost efficiency.
  • Design data mixtures, augmentations, and curriculum strategies for large robot datasets.
  • Develop data loaders and preprocessing for synchronized video, language, and robot state data.
  • Create reproducible experimentation systems with versioning and tracking.
  • Define offline and on-robot metrics for success, generalization, and safety.
  • Build tools to inspect trajectories and diagnose training issues.
  • Run experiments on physical robots to guide model and infrastructure improvements.
  • Collaborate with data-collection teams to improve coverage and data quality.
  • Translate research into production-quality systems for repeated deployment across robots.
  • Make technical decisions across architecture, data, and compute, communicating tradeoffs.
  • Mentor engineers and establish engineering practices for ML codebase.

Skills

Deep learning
PyTorch/JAX
Transformer models
Distributed training
Multimodal models
Experiment tracking
Communication skills
Debugging

Education

Bachelor's degree in CS/ML/Robotics

Tools

ROS/ROS 2
MuJoCo
Gazebo
PyTorch

Job description

About Us

Nidus is building autonomous manufacturing systems powered by AI and robotics. U.S. manufacturers are under growing pressure to increase output while navigating workforce constraints, underutilized equipment, and increasingly complex supply chains. We believe the next generation of manufacturing will be enabled by more intelligent, flexible automation-systems that help people and machines work together to expand capacity and improve productivity.

We're building the intelligence layer that makes that possible.

Our team comes from the intersection of defense technology and AI. We've built and shipped products to DoD customers, and we bring deep experience across ML, robotics, product development, and defense.

We're a small, in-person team in New York City working on hard problems that matter. If you want to build the future of American manufacturing, we'd love to hear from you.

About the role

We are building intelligent robots that can perceive the world, understand instructions, and perform useful physical tasks in dynamic, real-world environments.

We are looking for a Staff Machine Learning Engineer to develop and scale the vision-language-action models that power our robots. You will own major parts of the model development lifecycle—from dataset construction and training recipes to distributed training infrastructure, evaluation, and on-robot validation.

This is a hands-on role at the intersection of machine learning research, large-scale systems, and robotics. You will work closely with researchers, robotics engineers, data teams, and operators to turn new modeling ideas into reliable robot capabilities. The right person is equally comfortable investigating why a policy failed on a robot, designing the next training experiment, and improving the infrastructure required to run that experiment efficiently at scale.

What you’ll do
  • Train and improve vision-language-action models for robotic manipulation and other embodied tasks.

  • Develop model architectures and training recipes spanning imitation learning, behavior cloning, transformer- and diffusion-based policies, multimodal pretraining, and post-training.

  • Build scalable pipelines for pretraining, fine-tuning, evaluation, checkpointing, and model release.

  • Train models across multi-node GPU clusters while improving utilization, throughput, stability, and cost efficiency.

  • Design data mixtures, sampling strategies, augmentations, and curriculum approaches for large, heterogeneous robot datasets.

  • Develop data loaders and preprocessing systems for synchronized video, language, robot state, actions, and other sensor modalities.

  • Create reproducible experimentation systems, including configuration management, dataset and model versioning, experiment tracking, and automated regression testing.

  • Define offline and on-robot metrics that measure task success, generalization, robustness, latency, safety, and failure modes.

  • Build tools for inspecting trajectories, visualizing model behavior, comparing experiments, and diagnosing data or training failures.

  • Run structured experiments on physical robots and use the results to guide model, data, and infrastructure improvements.

  • Partner with data collection teams to identify coverage gaps, improve demonstration quality, and prioritize collection based on model performance.

  • Translate promising research into maintainable, production-quality systems that can support repeated deployment across robots and environments.

  • Make technical decisions across model architecture, data, compute, and evaluation, and communicate the associated tradeoffs clearly.

  • Mentor other engineers and establish strong engineering practices for the ML codebase.

What we’re looking for
  • Five or more years of professional experience in machine learning, robotics, computer vision, or a closely related field, or equivalent demonstrated impact.

  • Strong experience training modern deep learning models using PyTorch, JAX, or a comparable framework.

  • Experience with transformer-based models, multimodal models, generative models, or learned control policies.

  • Experience building and operating distributed training pipelines on multi-GPU or multi-node accelerator clusters.

  • Experience with large-scale datasets, high-throughput data loading, experiment tracking, and reproducible ML workflows.

  • Strong understanding of optimization, model architecture, data quality, evaluation methodology, and experimental design.

  • Ability to debug failures across the full stack, from input data and training dynamics to inference and robot behavior.

  • Experience independently owning technically ambiguous projects and delivering working systems.

  • Clear communication skills and an ability to collaborate across research, infrastructure, robotics, hardware, and operations teams.

  • A bachelor’s degree in computer science, machine learning, robotics, electrical engineering, or a related field—or equivalent practical experience.

Nice to have
  • Experience developing vision-language-action models, vision-language models, or robotics foundation models.

  • Experience with imitation learning, reinforcement learning, behavior cloning, diffusion policies, action tokenization, or action-conditioned world models.

  • Experience training policies using data from multiple robot embodiments, task domains, or sensor configurations.

  • Familiarity with large-scale pretraining, post-training, parameter-efficient fine-tuning, distillation, or model adaptation.

  • Experience optimizing distributed training using techniques such as FSDP, tensor or pipeline parallelism, mixed precision, gradient checkpointing, or sharded data loading.

  • Familiarity with robotics middleware and simulation environments such as ROS/ROS 2, MuJoCo, Isaac Sim, or Gazebo.

  • Experience deploying and evaluating learned policies on physical robotic systems.

  • Publications or meaningful open-source contributions in machine learning, robotics, computer vision, or related areas.


Nidus Technologies is an equal opportunity employer, and qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender perception or identity, national origin, age, marital status, protected veteran status, or disability status.

We encourage candidates from all backgrounds to apply, even if you don't feel like you're a perfect fit. If you're passionate about contributing to our mission, we'd love to hear from you!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Hardware Engineer, Robotics Systems & Deployments
Staff Hardware Engineer, Robotics Systems & Deployments

Nidus Technologies • New York (NY)

On-site
USD 150,000 - 230,000
Staff ML Engineer — Vision-Language for Robotics
Staff ML Engineer — Vision-Language for Robotics

Nidus • Northern (KY), New York (NY)

Hybrid
USD 120,000 - 180,000
Founding ML Engineer
Founding ML Engineer

a16z-speedrun • San Francisco (CA)

On-site
USD 150,000 - 230,000
Robot Learning Engineer
Robot Learning Engineer

Maxwell Bond • New York (NY)

On-site
USD 120,000 - 180,000
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda • Mountain View (CA)

On-site
USD 180,000 - 240,000
Deep Learning - Robot Manipulation Engineer
Deep Learning - Robot Manipulation Engineer

Persona AI, Inc. • Houston (TX)

On-site
USD 120,000 - 210,000
Competitive salary
Performance bonus
Medical benefits
+3
AI Engineer – Robotics Data Preprocessing
AI Engineer – Robotics Data Preprocessing

Persona AI Inc • Houston (TX)

On-site
USD 120,000 - 180,000
Performance-based bonus
Medical benefits 99% covered
Equity
+2
Research Member of Technical Staff- Robot Learning Systems & Reliability
Research Member of Technical Staff- Robot Learning Systems & Reliability

Rhoda AI • Mountain View (CA)

On-site
USD 190,000 - 270,000
Senior Robotics Engineer - R&D
Senior Robotics Engineer - R&D

Sereact • Boston (MA)

On-site
USD 150,000 - 225,000
Medical, dental, and vision insurance
401(k) with 100% company match, up to
20 days paid time off
+5
AI Engineer – Robotics Data Preprocessing
AI Engineer – Robotics Data Preprocessing

Persona AI, Inc. • Pensacola (FL), Northern (KY)

Hybrid
USD 120,000 - 160,000
Bonus
Medical benefits
Equity
+2