ML Infra Engineer: Scale Large-Scale Training Systems

physicalintelligence

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Physical Intelligence is seeking an hands-on ML Infrastructure Engineer to scale and optimize our training systems and core model code. You will own critical infrastructure for large-scale training, managing GPU/TPU compute, job orchestration, and building efficient JAX training pipelines.

You’ll collaborate with researchers and model engineers to translate ideas into experiments and production runs, contributing to core training code for new architectures and multimodal models.

Qualifications

  • Hands-on experience building ML training infrastructure.
  • Proficiency with JAX and PyTorch.
  • Familiarity with distributed training and multi-host setups.
  • Experience with cloud platforms (GCP, AWS) and orchestration tools.

Responsibilities

  • Own training/inference infrastructure including scheduling, checkpointing, and metrics.
  • Scale distributed training across TPU and GPU clusters.
  • Profile and optimize memory, throughput, and synchronization.
  • Build abstractions for launching, monitoring, and reproducing experiments.
  • Collaborate with researchers to translate ideas into production training runs.

Skills

Software engineering
Large-scale training
JAX
PyTorch
Distributed training
Cloud platforms

Tools

JAX
PyTorch
Kubernetes
SLURM
GCP TPU/GKE
AWS

Job description

Physical Intelligence is seeking an hands-on ML Infrastructure Engineer to scale and optimize our training systems and core model code. You will own critical infrastructure for large-scale training, managing GPU/TPU compute, job orchestration, and building efficient JAX training pipelines.

You’ll collaborate with researchers and model engineers to translate ideas into experiments and production runs, contributing to core training code for new architectures and multimodal models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

physicalintelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Infrastructure Engineer (Modeling)
Machine Learning Infrastructure Engineer (Modeling)

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Data Infrastructure Engineer
ML Data Infrastructure Engineer

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure Engineer: Training & Inference
ML Infrastructure Engineer: Training & Inference

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000
ML Infra Engineer — Scalable Training Systems
ML Infra Engineer — Scalable Training Systems

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000