ML Infrastructure Engineer: Training & Inference, Equity

Physical Superintelligence

Boston (MA)

Hybrid

USD 140,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Physical Superintelligence in Boston is hiring an ML infrastructure engineer to own and evolve the training and inference stack powering our scale-out research. You will design, implement, and operate distributed multi-GPU training jobs, model-serving pipelines, and the tooling researchers rely on to run experiments efficiently.

You will collaborate with engineering on capacity planning, observability, and reliability, and you will be the go-to person when pipelines stall or workloads fail.

Qualifications

  • Three or more years building and operating ML training or inference infrastructure in production.
  • Hands-on experience with distributed training (multi-GPU or multi-node) and model-serving systems.
  • Strong software engineering fundamentals for building reliable services.
  • ML fluency to work productively with researchers even if you don't design algorithms.

Responsibilities

  • Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems for both proprietary models and self-hosted inference.
  • Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models.
  • Partner with Engineering on capacity planning, observability, and reliability for GPU and inference infrastructure.
  • Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close.
  • Stay hands-on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.

Skills

Distributed training
PyTorch
Ray
Model-serving

Tools

vLLM
SGLang
Triton
CUDA

Job description

Physical Superintelligence in Boston is hiring an ML infrastructure engineer to own and evolve the training and inference stack powering our scale-out research. You will design, implement, and operate distributed multi-GPU training jobs, model-serving pipelines, and the tooling researchers rely on to run experiments efficiently.

You will collaborate with engineering on capacity planning, observability, and reliability, and you will be the go-to person when pipelines stall or workloads fail.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Physical Superintelligence • Boston (MA)

On-site
USD 140,000 - 230,000
Staff Cloud Platform Engineer - ML Infra & Release Automation
Staff Cloud Platform Engineer - ML Infra & Release Automation

Physical Superintelligence • Boston (MA)

Hybrid
USD 120,000 - 160,000
ML Infra Engineer: Scale Distributed Training & HPC
ML Infra Engineer: Scale Distributed Training & HPC

Watney • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Data Infrastructure Engineer
ML Data Infrastructure Engineer

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Performance Engineer: Scale GPU-Driven Training
ML Performance Engineer: Scale GPU-Driven Training

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer (Data Systems)
ML Infra Engineer (Data Systems)

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000