Research Engineer: Scalable ML Training Systems

Mind Robotics Inc.

Palo Alto, Northern (CA, KY)

Hybrid

USD 180,000 - 300,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mind Robotics is seeking a Research Engineer to build core systems for scalable ML model training on real-world industrial data. You will design distributed training pipelines across hundreds of GPUs, enabling fast experimentation and cost-efficient iterations.

Collaborate with modeling teams to accelerate iteration speed, reduce training costs, and develop internal tooling for experiment tracking, monitoring, and deployment.

Qualifications

  • Extensive experience building infrastructure for large-scale ML training.
  • Experience scaling distributed training across hundreds of GPUs.
  • Strong Python, PyTorch/JAX, and optimizer knowledge.

Responsibilities

  • Design and implement scalable systems for training large ML models.
  • Enable efficient workflows for data ingestion, training, and iteration.
  • Develop and optimize distributed training systems across hundreds of GPUs.
  • Implement strategies for parallelization, sharding, and efficient compute utilization.
  • Improve training efficiency through attention optimization, kernel fusion, and memory management.
  • Partner closely with modeling teams to accelerate iteration speed and reduce training costs.
  • Build internal tools for experiment tracking, monitoring, and debugging.
  • Implement systems for tracking training performance, failures, and resource utilization.
  • Debug and resolve bottlenecks across the training stack.
  • Provide lightweight infrastructure support for deploying and running models in production environments.
  • Optimize inference performance and reliability where needed.
  • Support core cloud infrastructure needs for training workloads (without heavy DevOps overhead).
  • Manage compute resources efficiently across training jobs.

Skills

Python
PyTorch
JAX
Distributed training
GPU scaling
Memory optimization
Research collaboration

Tools

Linux
Bash scripting
Experiment tracking

Job description

Mind Robotics is seeking a Research Engineer to build core systems for scalable ML model training on real-world industrial data. You will design distributed training pipelines across hundreds of GPUs, enabling fast experimentation and cost-efficient iterations.

Collaborate with modeling teams to accelerate iteration speed, reduce training costs, and develop internal tooling for experiment tracking, monitoring, and deployment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Distributed ML Training Systems Architect
Distributed ML Training Systems Architect

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Model Systems Engineer
Model Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer
Research Engineer

Mind Robotics Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Robotics AI Research & Modeling Engineer
Robotics AI Research & Modeling Engineer

Mind Robotics Inc. • Palo Alto (CA)

On-site
USD 100,000 - 130,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Research + Modeling
Research + Modeling

Mind Robotics Inc. • Palo Alto (CA)

On-site
USD 100,000 - 130,000
Staff ML Infra Engineer: Scalable Training & Pipelines
Staff ML Infra Engineer: Scalable Training & Pipelines

Watney Robotics Inc • San Francisco (CA)

On-site
USD 140,000 - 210,000
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
ML Research Engineer — Large-Scale Training & Tools
ML Research Engineer — Large-Scale Training & Tools

The Resume Database • Palo Alto (CA)

On-site
USD 120,000 - 160,000