Lead Machine Learning Engineer Serve Robotics · Remote · US · Machine Learning Engineering $225,000–$260,000 6mo ago

Aimlroles

Los Angeles, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Serve Robotics is seeking a senior ML engineer to design and scale large-scale training systems for multimodal robotics data, enabling high-performance autonomy models. You will optimize distributed pipelines, neural architectures, and data processing to improve training efficiency and GPU utilization.

The role collaborates with ML researchers and infrastructure teams to deploy models in production and guide model evolution through experiments and metrics.

Qualifications

  • Master’s or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or related field.
  • 5+ years of professional experience developing, training, and deploying ML models in production.
  • Hands-on experience training ML models across multiple GPUs with distributed training frameworks and large datasets.
  • Strong programming skills in Python for ML pipelines and training workflows.
  • Solid knowledge of neural networks, optimization algorithms, loss functions, model evaluation, and training methodologies.

Responsibilities

  • Design and maintain training systems that can process and learn from petabyte-scale multimodal datasets (e.g., video and point cloud data).
  • Identify and resolve bottlenecks in the training pipeline to maximize GPU utilization and reduce training time.
  • Work with the ML team to develop and refine neural network architectures for autonomy tasks.
  • Create and adjust loss functions and training strategies to improve autonomy performance.
  • Configure, monitor, and maintain large-scale distributed training jobs across multiple machines and GPUs.
  • Implement scalable systems to preprocess, transform, and augment large robotics datasets for model training.
  • Collaborate with ML scientists and engineers to integrate new models into production training pipelines.
  • Analyze training metrics and experiment logs to guide improvements in architecture and data usage.
  • Develop tools and workflows to run experiments, track results, and iterate on new model ideas.

Skills

Python programming
Multi-GPU training
Distributed training
Neural networks
Optimization algorithms
Training methodologies

Education

Master’s or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or related

Job description

At Serve Robotics, we’re reimagining how things move in cities. Our personable sidewalk robot is our vision for the future. It’s designed to take deliveries away from congested streets, make deliveries available to more people, and benefit local businesses.

The Serve fleet has been delighting merchants, customers, and pedestrians along the way in Los Angeles, Miami, Dallas, Atlanta and Chicago while doing commercial deliveries. We’re looking for talented individuals who will grow robotic deliveries from surprising novelty to efficient ubiquity.

Who We Are

We are tech industry veterans in software, hardware, and design who are pooling our skills to build the future we want to live in. We are solving real-world problems leveraging robotics, machine learning and computer vision, among other disciplines, with a mindful eye towards the end-to-end user experience. Our team is agile, diverse, and driven. We believe that the best way to solve complicated dynamic problems is collaboratively and respectfully.

This role develops and scales large-scale machine learning training systems for multimodal robotics data, enabling the creation of high-performance autonomy models. By optimizing distributed training pipelines, neural network architectures, and data processing workflows, the position improves training efficiency, accelerates model iteration, and maximizes GPU utilization. The role collaborates closely with ML researchers and infrastructure teams, influencing the design, deployment, and performance of end-to-end autonomy models and the large-scale data pipelines that support them.

Responsibilities

  • Design and maintain training systems that can process and learn from petabyte-scale multimodal datasets (e.g., video and point cloud data). This includes ensuring data is efficiently loaded, distributed, and processed across large GPU clusters.

  • Identify and resolve bottlenecks in the training pipeline, including data loading, preprocessing, model computation, and inter-node communication, to maximize GPU utilization and reduce training time.

  • Work with the ML team to develop and refine neural network architectures suitable for autonomy tasks, particularly those handling high-dimensional and sequential sensor data.

  • Create and adjust loss functions and training strategies that help the model learn effectively from complex multimodal inputs and improve autonomy performance.

  • Configure, monitor, and maintain large-scale distributed training jobs across multiple machines and GPUs, ensuring stability, fault tolerance, and efficient resource usage.

  • Implement scalable systems to preprocess, transform, and augment large robotics datasets so that they are suitable for model training.

  • Work closely with ML scientists and other engineers to integrate new models, experiments, and training approaches into the production training pipeline.

  • Analyze training metrics, model outputs, and experiment logs to assess model performance and guide improvements in architecture, data usage, or training strategies.

  • Develop tools and workflows that allow teams to run experiments, track results, and iterate quickly on new model ideas or training approaches.

Qualifications

  • Master’s or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a closely related technical discipline.

  • Minimum of 5 years of professional experience developing, training, and deploying machine learning models in production environments.

  • Hands‑on experience training machine learning models across multiple GPUs or compute nodes, including familiarity with distributed training frameworks and large dataset handling.

  • Strong programming skills in Python for implementing machine learning models, data pipelines, and training workflows.

  • Solid knowledge of core concepts such as neural networks, optimization algorithms, loss functions, model evaluation, and training methodologies.

What Makes You Stand out

  • Experience identifying and resolving training bottlenecks related to compute utilization, memory usage, and data throughput in machine learning systems.

  • Experience training machine learning models on robotics or autonomous driving datasets involving multimodal sensor inputs such as camera video, LiDAR point clouds, radar, or telemetry data.

  • Experience developing models that combine multiple data modalities (e.g., images, point clouds, and structured sensor data) into a unified learning system.

  • Peer‑reviewed publications or significant research contributions in machine learning, robotics, or related areas.

*Please note: The listed base salary range applies to candidates based in the US. Compensation may vary depending on location, experience, and role alignment. We are open to qualified candidates working remotely in Canada

  • Canada - ALL: $177k - $215k CAD

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Machine Learning Engineer
Lead Machine Learning Engineer

Devconnectplatform • United States

Hybrid
USD 225,000 - 260,000
Lead ML Engineer, Robotics Autonomy (Remote)
Lead ML Engineer, Robotics Autonomy (Remote)

Serve Robotics • Ottawa (IL)

On-site
USD 120,000 - 160,000
Lead ML Engineer, Robotics & Autonomy - Remote (Canada)
Lead ML Engineer, Robotics & Autonomy - Remote (Canada)

Aimlroles • Los Angeles (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Director of Product, Data Flywheel
Director of Product, Data Flywheel

Serve Robotics • Redwood City (CA)

On-site
USD 200,000 - 245,000
VP of IT Infrastructure & Operations at Serve Robotics
VP of IT Infrastructure & Operations at Serve Robotics

Matcha • Northern (KY)

Hybrid
USD 250,000 - 350,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

symbotic • Wilmington (MA)

On-site
USD 120,000 - 165,000
Medical benefits
Dental benefits
Vision benefits
+1
Lead ML Engineer - Robotics Autonomy (Remote)
Lead ML Engineer - Robotics Autonomy (Remote)

Devconnectplatform • United States

Hybrid
USD 225,000 - 260,000
ML Engineer - Robotics
ML Engineer - Robotics

Clera • United States

Remote
USD 220,000 - 300,000
Sr. Embedded Software Engineer, Robotics Platform
Sr. Embedded Software Engineer, Robotics Platform

Serve Robotics Inc • United States

Hybrid
USD 140,000 - 210,000
Systems Engineer - Machine Learning
Systems Engineer - Machine Learning

General Robotics • Redmond (WA)

On-site
USD 155,000 - 200,000
Medical benefits
401K