ML Infra Engineer: Scale Distributed Training & Research

Doist

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Doist is seeking an Infrastructure Research Engineer to own the core systems researchers rely on for distributed training, data pipelines, and experiment tooling. You will work directly with researchers, translating scientific needs into robust, scalable systems that operate across thousands of GPUs.

Strong Python and C++ skills, plus hands-on GPU profiling and memory optimization, are essential to accelerate research while maintaining reliability at scale.

Qualifications

  • Experience building and operating distributed training systems end-to-end.
  • Strong fundamentals in distributed systems, networking, storage, and performance analysis.
  • Proficiency in Python and C++ for research-oriented infrastructure.

Responsibilities

  • Own the core systems used by researchers for distributed training, experiment orchestration, and data pipelines.
  • Scale infrastructure to support thousands of GPUs and large-scale RL training.
  • Profile and optimize training throughput across data loading, compute, and memory.
  • Build and maintain tools for experiment tracking and reproducibility.
  • Diagnose and resolve failures across GPUs, networking, and numerics.

Skills

Distributed training systems
Python
C++
PyTorch
GPU profiling
Memory optimization
Parallelism
Debugging

Tools

PyTorch (systems level)
CUDA

Job description

Doist is seeking an Infrastructure Research Engineer to own the core systems researchers rely on for distributed training, data pipelines, and experiment tooling. You will work directly with researchers, translating scientific needs into robust, scalable systems that operate across thousands of GPUs.

Strong Python and C++ skills, plus hands-on GPU profiling and memory optimization, are essential to accelerate research while maintaining reliability at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Infrastructure Research Engineer - Distributed AI Training
Infrastructure Research Engineer - Distributed AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1