Machine Learning Infrastructure Engineer

Alexander Chapman

New York (NY)

On-site

USD 120,000 - 160,000

Full time

17 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Dental & Vision

Job summary

Alexander Chapman is seeking a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists rely on to train, deploy, and monitor models. You’ll focus on making systems around them fast, reliable, and easy to use.

Design and maintain training and inference infrastructure, build internal tooling for experiment tracking, feature stores, and model versioning, and optimize model serving for latency and cost at scale.

Qualifications

  • 2–5 years of experience in infrastructure, platform, or backend engineering, ideally supporting ML workloads.
  • Strong proficiency in Python and/or Go.
  • Experience with containerization and orchestration (Docker, Kubernetes).
  • Familiarity with ML-specific tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar.
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi).
  • Understanding of distributed systems and data pipeline design.
  • Comfort working with GPUs and understanding of training/serving performance trade-offs.
  • Strong communication skills and ability to work cross-functionally with ML practitioners.

Responsibilities

  • Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling).
  • Build internal tooling for experiment tracking, feature stores, and model versioning.
  • Optimize model serving for latency, throughput, and cost at scale.
  • Set up and maintain CI/CD pipelines for ML workflows.
  • Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm).
  • Implement monitoring and alerting for model performance, data drift, and system health.
  • Partner with ML engineers and data scientists to understand their workflows and remove friction.
  • Contribute to platform architecture decisions as the ML org scales.

Skills

Python
Go
Docker
Kubernetes
MLflow
Kubeflow
Ray
SageMaker
Vertex AI
Terraform
Pulumi

Tools

MLflow
Kubeflow
Ray
SageMaker
Vertex AI
Docker
Kubernetes
Terraform
Pulumi
Triton
TorchServe
vLLM
DVC
Feast
LakeFS

Job description

We're looking for a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists and ML engineers rely on to train, deploy, and monitor models. You'll focus less on the models themselves and more on making the systems around them fast, reliable, and easy to use.

What You'll Do
  • Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling)
  • Build internal tooling for experiment tracking, feature stores, and model versioning
  • Optimize model serving for latency, throughput, and cost at scale
  • Set up and maintain CI/CD pipelines for ML workflows
  • Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm)
  • Implement monitoring and alerting for model performance, data drift, and system health
  • Partner with ML engineers and data scientists to understand their workflows and remove friction
  • Contribute to platform architecture decisions as the ML org scales
What We're Looking For
  • 2–5 years of experience in infrastructure, platform, or backend engineering, ideally supporting ML workloads
  • Strong proficiency in Python and/or Go
  • Experience with containerization and orchestration (Docker, Kubernetes)
  • Familiarity with ML-specific tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi)
  • Understanding of distributed systems and data pipeline design
  • Comfort working with GPUs and understanding of training/serving performance trade-offs
  • Strong communication skills and ability to work cross-functionally with ML practitioners
Nice to Have
  • Experience with high-performance model serving (Triton, TorchServe, vLLM)
  • Familiarity with data versioning tools (DVC, LakeFS) or feature stores (Feast, Tecton)
  • Experience scaling distributed training (Horovod, DeepSpeed, PyTorch DDP)
  • Background in SRE or DevOps practices applied to ML systems
What We Offer
  • Competitive salary and equity
  • Health, dental, and vision insurance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior ML Infrastructure Engineer
Senior ML Infrastructure Engineer

Harnham • New York (NY)

On-site
USD 150,000 - 200,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
MLOps Engineer
MLOps Engineer

Evlo AI • Boston (MA)

On-site
USD 130,000 - 180,000
MLOps Engineer
MLOps Engineer

Mylitm • California (MO)

On-site
USD 140,000 - 210,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000