Machine Learning Infrastructure Engineer

Alexander Chapman Ltd

New York (NY)

On-site

USD 120,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Competitive equity

Job summary

Alexander Chapman Ltd is seeking a Mid-Level ML Infrastructure Engineer in New York to build and scale platforms and tooling for data scientists and ML engineers to train, deploy, and monitor models.

You will design pipelines, orchestrate compute, and implement CI/CD for ML workflows while collaborating with ML teams to reduce friction and improve performance and cost at scale.

Qualifications

  • 2–5 years of experience in infrastructure, platform, or backend engineering.
  • Proficient in Python and/or Go.
  • Experience with containerization and orchestration (Docker, Kubernetes).
  • Familiarity with ML tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar.
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi).
  • Understanding of distributed systems and data pipeline design; comfort with GPUs and training/serving performance trade-offs.
  • Strong communication skills and ability to work cross-functionally with ML practitioners.

Responsibilities

  • Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling).
  • Build internal tooling for experiment tracking, feature stores, and model versioning.
  • Optimize model serving for latency, throughput, and cost at scale.
  • Set up and maintain CI/CD pipelines for ML workflows.
  • Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm).
  • Implement monitoring and alerting for model performance, data drift, and system health.
  • Partner with ML engineers and data scientists to understand their workflows and remove friction.
  • Contribute to platform architecture decisions as the ML org scales.

Skills

Python
Go
Distributed systems
Cross-functional collaboration

Tools

Docker
Kubernetes
Terraform
Ray
MLflow
Kubeflow

Job description

We're looking for a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists and ML engineers rely on to train, deploy, and monitor models. You'll focus less on the models themselves and more on making the systems around them fast, reliable, and easy to use.

What You'll Do
  • Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling)
  • Build internal tooling for experiment tracking, feature stores, and model versioning
  • Optimize model serving for latency, throughput, and cost at scale
  • Set up and maintain CI/CD pipelines for ML workflows
  • Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm)
  • Implement monitoring and alerting for model performance, data drift, and system health
  • Partner with ML engineers and data scientists to understand their workflows and remove friction
  • Contribute to platform architecture decisions as the ML org scales
What We're Looking For
  • 2–5 years of experience in infrastructure, platform, or backend engineering, ideally supporting ML workloads
  • Strong proficiency in Python and/or Go
  • Experience with containerization and orchestration (Docker, Kubernetes)
  • Familiarity with ML-specific tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar
  • Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi)
  • Understanding of distributed systems and data pipeline designComfort working with GPUs and understanding of training/serving performance trade-offs
  • Strong communication skills and ability to work cross-functionally with ML practitioners
Nice to Have
  • Experience with high-performance model serving (Triton, TorchServe, vLLM)
  • Familiarity with data versioning tools (DVC, LakeFS) or feature stores (Feast, Tecton)
  • Experience scaling distributed training (Horovod, DeepSpeed, PyTorch DDP)
  • Background in SRE or DevOps practices applied to ML systems
What We Offer
  • Competitive salary and equity
  • Health, dental, and vision insurance
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Alexander Chapman • New York (NY)

On-site
USD 120,000 - 160,000
Equity
Health insurance
Dental & Vision
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

On-site
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
MLOps Engineer MLOps Engineer
MLOps Engineer MLOps Engineer

Kurai • Austin (TX)

On-site
USD 140,000 - 190,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Objective Partners • San Francisco (CA)

On-site
USD 180,000 - 250,000
Full medical, dental, vision coverage
Flexible PTO
Daily catered lunches
+1
ML Infra Engineer: Scale Pipelines, GPUs & Platforms (Equity)
ML Infra Engineer: Scale Pipelines, GPUs & Platforms (Equity)

Alexander Chapman • New York (NY)

On-site
USD 120,000 - 160,000
Equity
Health insurance
Dental & Vision
ML Infra Engineer, Modeling
ML Infra Engineer, Modeling

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
Confidential
MLOps Engineer
MLOps Engineer

Evlo AI • Seattle (WA)

On-site
USD 130,000 - 190,000