Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor

Boston (MA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Core AI is seeking an experienced infrastructure engineer to own the training and inference platform. You will run distributed multi-GPU training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, Triton) across proprietary models and self-hosted inference.

You will build tooling researchers rely on to launch training runs, route workloads, and ensure capacity, observability, and reliability. The role involves debugging under real load and working hands-on, turning research code into

Qualifications

  • Three or more years building and operating ML training or inference infrastructure in production.
  • Hands-on experience with distributed training (multi-GPU or multi-node) and model-serving systems.
  • Strong software engineering fundamentals for reliable, scalable services.
  • ML fluency to debug training loops and inference-time behavior with researchers.

Responsibilities

  • Own the training and inference infrastructure for Core AI, including distributed training jobs, GPU scheduling, and model-serving systems.
  • Develop tools and abstractions for researchers to launch training runs and route workloads across models.
  • Partner with Engineering on capacity planning, observability, and reliability for GPU and inference infrastructure.
  • Debug and harden the training and inference stack under real load.
  • Stay hands-on by writing code and leading when a training job stalls or an inference path breaks.

Tools

PyTorch
Ray
vLLM
SGLang
Triton
Terraform
CUDA
GCP
AWS

Job description

Core AI is seeking an experienced infrastructure engineer to own the training and inference platform. You will run distributed multi-GPU training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, Triton) across proprietary models and self-hosted inference.

You will build tooling researchers rely on to launch training runs, route workloads, and ensure capacity, observability, and reliability. The role involves debugging under real load and working hands-on, turning research code into

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer: Training & Inference, Equity
ML Infrastructure Engineer: Training & Inference, Equity

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 230,000
Senior ML Platform Engineer - Scalable GPU AI Infra
Senior ML Platform Engineer - Scalable GPU AI Infra

Adobe Inc. • San Jose (CA)

On-site
USD 183,000 - 265,000
ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior ML Infra Engineer for Distributed GPU Training
Senior ML Infra Engineer for Distributed GPU Training

Genesis Molecular AI • City of Utica (NY)

On-site
USD 150,000 - 190,000
Competitive compensation with salary +
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Infra Engineer - Large-Scale Training & Pipelines
Senior ML Infra Engineer - Large-Scale Training & Pipelines

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays