ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Physical Intelligence in San Francisco is seeking a hands-on ML infrastructure engineer to scale and optimize our training systems. You will own large-scale training infra, manage GPU/TPU compute, job orchestration, checkpointing, and memory profiling, building efficient JAX pipelines that turn research ideas into production runs.

You’ll work with researchers and model engineers to push the limits of scalable ML, contribute to core training code, and help ensure reliability and speed across

Qualifications

  • Strong software engineering fundamentals and ML infra experience.
  • Hands-on large-scale training experience in JAX or PyTorch.
  • Experience with distributed training and multi-host setups.
  • Experience deploying training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS).
  • Ability to debug and optimize performance bottlenecks in the training stack.

Responsibilities

  • Own training/inference infrastructure: design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.
  • Scale distributed training: work with researchers to scale JAX-based training across TPU and GPU clusters.
  • Optimize performance: profile and improve memory usage, device utilization, throughput, and distributed synchronization.
  • Enable rapid iteration: build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • Partner with researchers: translate research needs into infra capabilities and guide best practices for training at scale.
  • Contribute to core training code: evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.

Skills

Software engineering fundamentals
Large-scale ML training
JAX
PyTorch
Distributed training
Cloud platforms
Performance debugging
Cross-functional collaboration

Tools

SLURM
Kubernetes
GCP TPU/GKE
AWS

Job description

Physical Intelligence in San Francisco is seeking a hands-on ML infrastructure engineer to scale and optimize our training systems. You will own large-scale training infra, manage GPU/TPU compute, job orchestration, checkpointing, and memory profiling, building efficient JAX pipelines that turn research ideas into production runs.

You’ll work with researchers and model engineers to push the limits of scalable ML, contribute to core training code, and help ensure reliability and speed across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer for Petabyte-Scale Training (SF On-Site)
ML Infra Engineer for Petabyte-Scale Training (SF On-Site)

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 400,000
ML Infra Engineer — Scalable Training Systems
ML Infra Engineer — Scalable Training Systems

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer (Data Systems)
ML Infra Engineer (Data Systems)

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000