Post-training ML Engineer

Autolab

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Autolab is building the orchestration and storage backbone for thousands of ML experiments in a production-grade environment. You will design and operate the scheduling layer, optimize GPU utilization, and own the artifact storage and observability for large-scale workflows.

You’ll work with Python and systems languages (Go, Rust, or C++) and gain hands-on experience with Kubernetes, Slurm, and Ray to support diverse GPU fleets and HPC workloads.

Qualifications

  • Experience building distributed systems in production: schedulers, queues, or storage.
  • Strong Python plus a systems language (Go, Rust, or C++).
  • Hands-on experience with cluster tooling such as Kubernetes, Slurm, or Ray.

Responsibilities

  • Design and operate the orchestration layer that schedules thousands of concurrent experiments across heterogeneous GPU fleets.
  • Build GPU scheduling that keeps clusters near full utilization, from a single lab 3090 to multi-node H100 clusters.
  • Own artifact storage for checkpoints, datasets, and logs: versioned, deduplicated, and fast to fetch.
  • Build the observability pipeline that ingests training telemetry at high volume without slowing runs down.
  • Make experiment launches reproducible: environments, dependencies, and data pinned by default.
  • Set the engineering practices, CI, deployment, on-call, that the team grows into.

Skills

Python
Go
Rust
C++

Tools

Kubernetes
Slurm
Ray

Job description

Build the distributed systems that run thousands of concurrent ML experiments: job orchestration, GPU scheduling, artifact storage, observability.

What You’ll Do
  • Design and operate the orchestration layer that schedules thousands of concurrent experiments across heterogeneous GPU fleets.
  • Build GPU scheduling that keeps clusters near full utilization, from a single lab 3090 to multi-node H100 clusters.
  • Own artifact storage for checkpoints, datasets, and logs: versioned, deduplicated, and fast to fetch.
  • Build the observability pipeline that ingests training telemetry at high volume without slowing runs down.
  • Make experiment launches reproducible: environments, dependencies, and data pinned by default.
  • Set the engineering practices, CI, deployment, on-call, that the team grows into.
What We’re Looking For
  • You have built and operated distributed systems in production: schedulers, queues, or storage.
  • Strong Python plus a systems language (Go, Rust, or C++).
  • Hands-on experience with cluster tooling such as Kubernetes, Slurm, or Ray.
  • You have run GPU or HPC workloads and know where they break.
  • Pragmatic by default: ship the simple version, measure, then harden.
  • Comfortable owning large areas with little process in an early-stage team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: GPU Orchestration & Observability
ML Infra Engineer: GPU Orchestration & Observability

Autolab • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer, ML Infrastructure
Software Engineer, ML Infrastructure

Cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Founding Engineer - ML Platforms
Founding Engineer - ML Platforms

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • San Francisco (CA)

On-site
USD 120,000 - 150,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Principal AI Engineer
Senior Principal AI Engineer

cerence • United States

On-site
USD 180,000 - 250,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000