ML Infrastructure Engineer: Scalable Training and Deployment

Epsilon

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Epsilon Health in San Francisco is seeking an ML infrastructure engineer to design core systems for scalable training of large models in medical imaging. You will enable researchers to run experiments efficiently, focusing on science rather than bottlenecks.

In this role you own distributed training, data loading, inference and deployment pipelines, and the RL training stack; collaborate with research and backend teams to ship production-grade solutions.

Qualifications

  • 6+ years designing, building, and operating large-scale distributed systems in production.
  • 2+ years building ML infrastructure or systems in production.
  • Strong Python skills and expertise in PyTorch or JAX.
  • Experience and familiarity with the compute, tooling, and workflow needs of large-scale ML research.
  • Experience building infrastructure or platforms specifically for research or ML workflows.
  • Deep experience building and operating Kubernetes and cloud infrastructure at scale.
  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism).
  • Prior experience as a technical lead or mentor for other engineers.

Responsibilities

  • Partner directly with researchers to understand workflows and design needs.
  • Build a distributed training infrastructure for foundation models on large-scale medical imaging.
  • Build high-throughput data loading and preprocessing to keep GPUs saturated.
  • Prototype ideas and translate them into production-ready code with end-to-end delivery.
  • Contribute to production serving and deployment pipelines alongside the backend team.
  • Build the reinforcement learning training stack for online, multi-reward RL at scale.

Skills

Distributed systems
ML infrastructure
Python
PyTorch
JAX
Kubernetes
Cloud infrastructure
Large-scale training
Leadership
Research workflows

Tools

DeepSpeed
Megatron
TensorRT
Triton
vLLM

Job description

Epsilon Health in San Francisco is seeking an ML infrastructure engineer to design core systems for scalable training of large models in medical imaging. You will enable researchers to run experiments efficiently, focusing on science rather than bottlenecks.

In this role you own distributed training, data loading, inference and deployment pipelines, and the RL training stack; collaborate with research and backend teams to ship production-grade solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer for Scalable Medical Imaging AI
ML Infrastructure Engineer for Scalable Medical Imaging AI

Epsilon Health • San Francisco (CA)

On-site
USD 150,000 - 250,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer - ML Infrastructure
Software Engineer - ML Infrastructure

Epsilon Health • San Francisco (CA)

On-site
USD 150,000 - 250,000
ML Infrastructure Engineer: Scale & Performance
ML Infrastructure Engineer: Scale & Performance

Physical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Production ML Engineer - Systems & Reliability
Production ML Engineer - Systems & Reliability

Sprinter Health • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Medical, dental, and vision plans 100%
Flexible PTO
401(k) with match
+2
ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
ML Infra Engineer: Scale, GPU Performance & Reliability
ML Infra Engineer: Scale, GPU Performance & Reliability

Jobtailor • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Senior Data Infrastructure Engineer – Health AI
Senior Data Infrastructure Engineer – Health AI

Epsilon Health • San Francisco (CA)

On-site
USD 150,000 - 250,000