ML Infra Engineer: Scale Pipelines, GPUs & Platforms (Equity)

Alexander Chapman

New York (NY)

On-site

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Dental & Vision

Job summary

Alexander Chapman is seeking a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists rely on to train, deploy, and monitor models. You’ll focus on making systems around them fast, reliable, and easy to use.

Design and maintain training and inference infrastructure, build internal tooling for experiment tracking, feature stores, and model versioning, and optimize model serving for latency and cost at scale.

Qualifications

  • 2–5 years of experience in infrastructure, platform, or backend engineering, ideally supporting ML workloads.
  • Strong proficiency in Python and/or Go.
  • Experience with containerization and orchestration (Docker, Kubernetes).
  • Familiarity with ML-specific tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar.
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi).
  • Understanding of distributed systems and data pipeline design.
  • Comfort working with GPUs and understanding of training/serving performance trade-offs.
  • Strong communication skills and ability to work cross-functionally with ML practitioners.

Responsibilities

  • Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling).
  • Build internal tooling for experiment tracking, feature stores, and model versioning.
  • Optimize model serving for latency, throughput, and cost at scale.
  • Set up and maintain CI/CD pipelines for ML workflows.
  • Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm).
  • Implement monitoring and alerting for model performance, data drift, and system health.
  • Partner with ML engineers and data scientists to understand their workflows and remove friction.
  • Contribute to platform architecture decisions as the ML org scales.

Skills

Python
Go
Docker
Kubernetes
MLflow
Kubeflow
Ray
SageMaker
Vertex AI
Terraform
Pulumi

Tools

MLflow
Kubeflow
Ray
SageMaker
Vertex AI
Docker
Kubernetes
Terraform
Pulumi
Triton
TorchServe
vLLM
DVC
Feast
LakeFS

Job description

Alexander Chapman is seeking a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists rely on to train, deploy, and monitor models. You’ll focus on making systems around them fast, reliable, and easy to use.

Design and maintain training and inference infrastructure, build internal tooling for experiment tracking, feature stores, and model versioning, and optimize model serving for latency and cost at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale RL Pipelines & Inference
ML Infra Engineer: Scale RL Pipelines & Inference

Moonfire • Paris (TX)

On-site
USD 115,000 - 173,000
Equity
Flexible time off
Relocation package
+4
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Alexander Chapman • New York (NY)

On-site
USD 120,000 - 160,000
Equity
Health insurance
Dental & Vision
ML Infra Engineer — Scale GPU ML Platform & Equity
ML Infra Engineer — Scale GPU ML Platform & Equity

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Senior ML Platform & Infra Engineer - Scale AI Pipelines
Senior ML Platform & Infra Engineer - Scale AI Pipelines

Monograph • United States

Hybrid
USD 160,000 - 240,000
Competitive base pay
Equity (RSUs)
Benefits
ML Infra Platform Engineer – Scale AI Accelerators
ML Infra Platform Engineer – Scale AI Accelerators

Amazon.com Services LLC • Seattle (WA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Scale GPU Training & Data Pipelines
ML Infra Engineer: Scale GPU Training & Data Pipelines

Humble Robotics • United States

On-site
USD 150,000 - 230,000
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Segment (Twilio) • San Francisco (CA)

On-site
USD 175,000 - 220,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2
MLOps Engineer — Scalable AI Infra & Deployment, Equity
MLOps Engineer — Scalable AI Infra & Deployment, Equity

Fundamental • United States

Remote
USD 180,000 - 260,000
Salary + equity
Health coverage for you and dependents
Parental leave for all
+2