ML Infrastructure Engineer — Build Scale Training & Serving

Harrison Clarke

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Harrison Clarke partners with a well-funded frontier AI lab to pursue a Machine Learning Infrastructure Engineer who will design and build the core infrastructure powering next-generation large-scale model training and deployment.

You will own the stack for training orchestration, distributed compute, and model serving, collaborating with researchers to deliver reliable, scalable platforms that accelerate AI research and production.

Qualifications

  • 5+ years of software engineering with ML infra focus.
  • Experience with ML tooling and distributed systems.
  • Expertise in Python and GPU programming.

Responsibilities

  • Design foundational ML infrastructure for training orchestration, distributed compute frameworks, GPU/TPU cluster management, and experiment tracking systems.
  • Optimize large-scale training pipelines for throughput, fault tolerance, checkpointing, and resource utilization across multi-node, multi-accelerator environments.
  • Develop and maintain model serving infrastructure with low-latency inference, packaging, deployment pipelines, and autoscaling.
  • Build internal platforms and tooling for self-service compute, data pipelines, and reproducible experiment workflows.
  • Collaborate closely with research teams to translate bleeding-edge ML requirements into reliable, scalable systems.
  • Drive infrastructure reliability with monitoring, observability, capacity planning, and incident response for mission-critical ML workloads.

Skills

Python
Distributed systems
ML infrastructure
GPUs/accelerators
Kubernetes
Large-scale training
Experiment tracking

Tools

PyTorch
JAX
DeepSpeed
Megatron
CUDA
NCCL
InfiniBand

Job description

Harrison Clarke partners with a well-funded frontier AI lab to pursue a Machine Learning Infrastructure Engineer who will design and build the core infrastructure powering next-generation large-scale model training and deployment.

You will own the stack for training orchestration, distributed compute, and model serving, collaborating with researchers to deliver reliable, scalable platforms that accelerate AI research and production.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer — Performance & Systems
AI Infrastructure Engineer — Performance & Systems

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML-Driven Infra Engineer: Scale Data Pipelines & Models
ML-Driven Infra Engineer: Scale Data Pipelines & Models

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff
Member of Technical Staff

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 280,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Platform Engineer - Scalable ML Infra
AI Platform Engineer - Scalable ML Infra

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000
ML Infrastructure Engineer — Data & Training
ML Infrastructure Engineer — Data & Training

Arena Physica • New York (NY)

On-site
USD 150,000 - 230,000
Premium medical, vision, dental
401(k) plan
Unlimited PTO
+2
ML Infrastructure Engineer - Scale & Deploy AI Models
ML Infrastructure Engineer - Scale & Deploy AI Models

SNAP, Inc. • Iowa (LA)

Hybrid
USD 150,000 - 230,000
Parental leave & fertility benefits
Health, dental, vision insurance
401(k) with company match
+1
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1