Member of Technical Staff, ML Engineer

Jobtailor

Boston (MA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Core AI is seeking an experienced infrastructure engineer to own the training and inference platform. You will run distributed multi-GPU training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, Triton) across proprietary models and self-hosted inference.

You will build tooling researchers rely on to launch training runs, route workloads, and ensure capacity, observability, and reliability. The role involves debugging under real load and working hands-on, turning research code into

Qualifications

  • Three or more years building and operating ML training or inference infrastructure in production.
  • Hands-on experience with distributed training (multi-GPU or multi-node) and model-serving systems.
  • Strong software engineering fundamentals for reliable, scalable services.
  • ML fluency to debug training loops and inference-time behavior with researchers.

Responsibilities

  • Own the training and inference infrastructure for Core AI, including distributed training jobs, GPU scheduling, and model-serving systems.
  • Develop tools and abstractions for researchers to launch training runs and route workloads across models.
  • Partner with Engineering on capacity planning, observability, and reliability for GPU and inference infrastructure.
  • Debug and harden the training and inference stack under real load.
  • Stay hands-on by writing code and leading when a training job stalls or an inference path breaks.

Tools

PyTorch
Ray
vLLM
SGLang
Triton
Terraform
CUDA
GCP
AWS

Job description

  • Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, or comparable) for both proprietary models and self-hosted inference.
  • Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher's time goes into the science instead of the plumbing.
  • Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else we ship.
  • Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close, not someone else's ticket.
  • Stay hands‑on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.
Requirements
  • Three or more years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.
  • Hands‑on experience with distributed training (multi-GPU or multi‑node, using PyTorch, Ray, or comparable) and model‑serving systems (vLLM, SGLang, Triton, or comparable).
  • Strong software engineering fundamentals. You can build a service that other engineers and researchers depend on every day, not a script that worked once.
  • Enough ML fluency to work productively with AI researchers: you understand training loops, reward signals, and inference‑time behavior well enough to debug them, even without designing the algorithms yourself.
  • Experience building internal platform tools such as training‑as‑a‑service APIs, inference gateways, or job schedulers.
  • Background in GPU infrastructure, CUDA, or performance engineering for ML workloads.
  • Experience with cloud infrastructure (GCP, AWS) and infrastructure as code (Terraform or comparable).
  • Prior work embedded alongside a research team, turning research code into production systems.
Core Competencies

Demonstrates expertise in building and operating machine learning training and inference infrastructure, with a strong focus on distributed training, model‑serving systems, and cloud infrastructure. Proficient in developing internal platform tools and collaborating effectively with AI researchers to enhance production systems.

Highest-signal resume keywords
  • Distributed Training (Multi-GPU, Multi-Node)
  • Model-Serving Systems (vLLM, SGLang, Triton)
  • Cloud Infrastructure (GCP, AWS)
  • Infrastructure as Code (Terraform)
  • GPU Infrastructure, CUDA
ATS Optimization Keywords
Hard Skills
  • Machine Learning Infrastructure
  • Software Engineering Fundamentals
  • Training Loops
  • Reward Signals
  • Inference-Time Behavior
  • Training-as-a-Service APIs
  • Inference Gateways
  • Job Schedulers
  • Performance Engineering
  • Debugging and Hardening Systems
Soft Skills
  • Collaboration
  • Problem-Solving
  • Hands-On Development
Industry Keywords
  • AI Research
  • Production Systems
  • Capacity Planning
  • Observability
  • Reliability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineering Technical Lead
AI Engineering Technical Lead

Jobtailor • Burbank (CA)

On-site
USD 180,000 - 240,000
AI and ML Infra Software Engineer, GPU Clusters
AI and ML Infra Software Engineer, GPU Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
AI Implementation Engineer
AI Implementation Engineer

Jobtailor • New Jersey

On-site
USD 150,000 - 210,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Director of Machine Learning
Director of Machine Learning

Jobtailor • Washington

On-site
USD 180,000 - 260,000
Senior AI Systems and Algorithms Engineer
Senior AI Systems and Algorithms Engineer

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Applied AI Scientist, Senior/Staff
Applied AI Scientist, Senior/Staff

Jobtailor • United States

On-site
USD 120,000 - 180,000
Distinguished Data Scientist
Distinguished Data Scientist

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000