ML Infra Engineer: GPU Orchestration & Observability

Autolab

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Autolab is building the orchestration and storage backbone for thousands of ML experiments in a production-grade environment. You will design and operate the scheduling layer, optimize GPU utilization, and own the artifact storage and observability for large-scale workflows.

You’ll work with Python and systems languages (Go, Rust, or C++) and gain hands-on experience with Kubernetes, Slurm, and Ray to support diverse GPU fleets and HPC workloads.

Qualifications

  • Experience building distributed systems in production: schedulers, queues, or storage.
  • Strong Python plus a systems language (Go, Rust, or C++).
  • Hands-on experience with cluster tooling such as Kubernetes, Slurm, or Ray.

Responsibilities

  • Design and operate the orchestration layer that schedules thousands of concurrent experiments across heterogeneous GPU fleets.
  • Build GPU scheduling that keeps clusters near full utilization, from a single lab 3090 to multi-node H100 clusters.
  • Own artifact storage for checkpoints, datasets, and logs: versioned, deduplicated, and fast to fetch.
  • Build the observability pipeline that ingests training telemetry at high volume without slowing runs down.
  • Make experiment launches reproducible: environments, dependencies, and data pinned by default.
  • Set the engineering practices, CI, deployment, on-call, that the team grows into.

Skills

Python
Go
Rust
C++

Tools

Kubernetes
Slurm
Ray

Job description

Autolab is building the orchestration and storage backbone for thousands of ML experiments in a production-grade environment. You will design and operate the scheduling layer, optimize GPU utilization, and own the artifact storage and observability for large-scale workflows.

You’ll work with Python and systems languages (Go, Rust, or C++) and gain hands-on experience with Kubernetes, Slurm, and Ray to support diverse GPU fleets and HPC workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Post-training ML Engineer
Post-training ML Engineer

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform ML Infra Engineer: GPU, Kubernetes & MLOps
Platform ML Infra Engineer: GPU, Kubernetes & MLOps

Oracle • United States

On-site
USD 92,000 - 210,000
Health insurance
401(k) match
Paid time off
+2
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
Founding ML Platforms Engineer: GPU Orchestration & Scale
Founding ML Platforms Engineer: GPU Orchestration & Scale

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options