Distributed ML Training Systems Architect

Mind Robotics

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mind Robotics is seeking a Model Systems Engineer to build core systems that enable fast, reliable, and scalable model training—from experimentation to production deployment.

You will design and implement scalable training systems, optimize distributed workflows across hundreds of GPUs, and collaborate closely with modeling teams to accelerate iteration speed while reducing training costs. Expect a rigorous, infrastructure-focused role in a cutting-edge robotics lab.

Qualifications

  • Experience building infrastructure for large-scale ML training.
  • Deep understanding of how modern LLM/VLM systems are trained and scaled.
  • Proven experience setting up and scaling distributed training across hundreds of GPUs.

Responsibilities

  • Design and implement scalable systems for training large ML models.
  • Enable efficient workflows for data ingestion, training, and iteration.
  • Develop and optimize distributed training systems across hundreds of GPUs.
  • Implement strategies for parallelization, sharding, and efficient compute utilization.
  • Improve training efficiency through attention optimization, kernel fusion, and memory management.
  • Partner with modeling teams to accelerate iteration speed and reduce training costs.
  • Build internal tools for experiment tracking, monitoring, and debugging.
  • Track training performance, failures, and resource utilization.
  • Debug and resolve bottlenecks across the training stack.
  • Provide lightweight infrastructure support for deploying and running models in production environments.
  • Optimize inference performance and reliability where needed.
  • Support core cloud infrastructure needs for training workloads.

Skills

Python
Distributed training
Memory management
Parallelism

Tools

PyTorch
JAX

Job description

Mind Robotics is seeking a Model Systems Engineer to build core systems that enable fast, reliable, and scalable model training—from experimentation to production deployment.

You will design and implement scalable training systems, optimize distributed workflows across hundreds of GPUs, and collaborate closely with modeling teams to accelerate iteration speed while reducing training costs. Expect a rigorous, infrastructure-focused role in a cutting-edge robotics lab.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed ML Training Systems Engineer
Distributed ML Training Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Model Systems Engineer
Model Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer
Research Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Robotics AI Research & Modeling Engineer
Robotics AI Research & Modeling Engineer

Mind Robotics Inc. • Palo Alto (CA)

On-site
USD 100,000 - 130,000
Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
Senior ML Training Systems Engineer, Large-Scale Robotics
Senior ML Training Systems Engineer, Large-Scale Robotics

Rhoda AI • Mountain View (CA)

On-site
USD 150,000 - 200,000
Distributed ML Engineer for Robotics & Multimodal AI
Distributed ML Engineer for Robotics & Multimodal AI

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Relocation assistance
Remote-Ready ML Systems Engineer: High-Throughput Training
Remote-Ready ML Systems Engineer: High-Throughput Training

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 160,000
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000