ML Infrastructure Engineer: Training & Inference

Physical Superintelligence

Boston (MA)

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Physical Superintelligence in Boston seeks a Member of Technical Staff, ML Engineer to build and run training and inference systems at scale, enabling researchers to focus on science rather than plumbing.

You will own distributed training, multi-GPU inference, and model-serving stacks (vLLM, SGLang, Triton), collaborate with engineers on observability and reliability, and ship with hands-on coding from day one.

Qualifications

  • Three or more years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.
  • Hands-on experience with distributed training (multi-GPU or multi-node, using PyTorch, Ray, or comparable) and model-serving systems (vLLM, SGLang, Triton, or comparable).
  • Strong software engineering fundamentals. You can build a service that other engineers and researchers depend on every day, not a script that worked once.
  • Enough ML fluency to work productively with AI researchers: you understand training loops, reward signals, and inference-time behavior well enough to debug them, even without designing the algorithms yourself.

Responsibilities

  • Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, or comparable) for both proprietary models and self-hosted inference.
  • Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher’s time goes into the science instead of the plumbing.
  • Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else we ship.
  • Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close, not someone else’s ticket.
  • Stay hands-on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.

Skills

Distributed training
Multi-GPU / Multi-node
PyTorch
Ray
Model serving
Software engineering fundamentals
Collaboration with researchers

Tools

vLLM
SGLang
Triton
CUDA
Terraform
AWS
GCP

Job description

Physical Superintelligence in Boston seeks a Member of Technical Staff, ML Engineer to build and run training and inference systems at scale, enabling researchers to focus on science rather than plumbing.

You will own distributed training, multi-GPU inference, and model-serving stacks (vLLM, SGLang, Triton), collaborate with engineers on observability and reliability, and ship with hands-on coding from day one.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000
ML Inference Infrastructure Architect
ML Inference Infrastructure Architect

Morph • California (MO)

On-site
USD 90,000 - 120,000
ML Data Infrastructure Engineer
ML Data Infrastructure Engineer

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Engineer - AI Inference & Benchmarking
Staff Engineer - AI Inference & Benchmarking

Liquid-Ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Equity
Health insurance
401(k) match
+1
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Cloud Platform Engineer - ML Infra & Release Automation
Staff Cloud Platform Engineer - ML Infra & Release Automation

Physical Superintelligence • Boston (MA)

Hybrid
USD 120,000 - 160,000
LLM Training & Inference Systems Engineer
LLM Training & Inference Systems Engineer

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 190,000 - 237,000