ML Infrastructure Engineer: Scale & Performance

Physical Intelligence

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Physical Intelligence in San Francisco is building the core ML infrastructure to scale training from prototype to production-grade runs. The ML Infrastructure team owns training/inference systems, scheduling, checkpointing, and metrics collection.

You will scale JAX-based training across TPU and GPU clusters, profile memory usage, improve throughput, and create abstractions for launching and monitoring experiments in close collaboration with researchers.

Responsibilities

  • Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging
  • Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction
  • Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization
  • Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments
  • Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale
  • Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics

Job description

Physical Intelligence in San Francisco is building the core ML infrastructure to scale training from prototype to production-grade runs. The ML Infrastructure team owns training/inference systems, scheduling, checkpointing, and metrics collection.

You will scale JAX-based training across TPU and GPU clusters, profile memory usage, improve throughput, and create abstractions for launching and monitoring experiments in close collaboration with researchers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale Large-Scale Training Systems
ML Infra Engineer: Scale Large-Scale Training Systems

physicalintelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

physicalintelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Infrastructure Engineer (Modeling)
Machine Learning Infrastructure Engineer (Modeling)

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infrastructure Engineer: Training & Inference
ML Infrastructure Engineer: Training & Inference

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
AI Performance Engineer: Scale ML Workloads at Datacenter
AI Performance Engineer: Scale ML Workloads at Datacenter

Applied Intuition • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Staff ML Compute & TPU Infrastructure Engineer
Staff ML Compute & TPU Infrastructure Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
ML Infrastructure Engineer for Scalable Medical Imaging AI
ML Infrastructure Engineer for Scalable Medical Imaging AI

Epsilon Health • San Francisco (CA)

On-site
USD 150,000 - 250,000