Software Engineer - ML Infrastructure

Epsilon Health

San Francisco (CA)

On-site

USD 150,000 - 250,000

Full time

29 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Epsilon Health is seeking an ML infrastructure engineer in San Francisco to design and build core systems enabling scalable training of large models for medical imaging and research. You will own distributed training, RL infrastructure, and inference pipelines from experimentation to production, collaborating with researchers and backend teams.

Responsibilities include partnering with researchers, building high-throughput data loading, and shipping end-to-end systems for training, evaluation,

Qualifications

  • 6+ years designing, building, and operating large-scale distributed systems.

Responsibilities

  • Partner with researchers to understand workflows and future needs.
  • Build distributed training infra for foundation models on large medical imaging data.
  • Develop high-throughput data loading to keep GPUs saturated.
  • Prototype ideas and ship production-ready code end-to-end.
  • Contribute to production serving, deployment pipelines, and monitoring.
  • Build RL training stack for online, multi-reward RL at scale.

Skills

Distributed systems
ML infrastructure
Python
PyTorch/JAX
Kubernetes
DeepSpeed/Megatron
Research workflows
Reinforcement learning

Tools

vLLM
TensorRT
Trition
Cloud infrastructure

Job description

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

Role Overview

We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks.

Sitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production.

Key Responsibilities
  • Partner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change
  • Build a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.
  • Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.
  • Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.
  • Contribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team.
  • Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.
Qualifications
  • 6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production
  • Have 2+ years of experience building ML infrastructure or systems in production
  • Strong Python skills and expertise in PyTorch or JAX
  • Experience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research
  • Experience building infrastructure or platforms specifically for research or machine learning workflows
  • Deep experience building and operating Kubernetes and cloud infrastructure at scale
  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient
  • Prior experience as a technical lead or mentor for other engineers
Preferred Qualifications
  • Experience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy
  • Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems
  • Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production
  • Experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.
  • Experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.
  • Experience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows
  • Familiarity with vision-language models (VLMs) or multimodal architectures

The anticipated annual base salary for this position is up to $250,000. This range does not include any other compensation components or other benefits for which an individual may be eligible. The actual base salary offered depends on a variety of factors, which may include as applicable, the qualifications of the individual applicant for the position, years of relevant experience, specific and unique skills, level of education attained, certifications or other professional licenses held, and the location in which the applicant lives and/or from which they will be performing the job.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Core Systems
Software Engineer - Core Systems

Epsilon Health • San Francisco (CA)

On-site
USD 150,000 - 250,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
ML Infrastructure Engineer
ML Infrastructure Engineer

echo inc. • San Francisco (CA)

On-site
USD 180,000 - 230,000
Stock options
Competitive compensation
401(k) program with matching
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Echo • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including stock options
Comprehensive benefits package
401(k) program with matching contributions
ML Infrastructure EngineerSan Francisco
ML Infrastructure EngineerSan Francisco

Stealth Neurotechnology Company • San Francisco (CA)

On-site
USD 180,000 - 230,000
Stock options
Comprehensive benefits
401(k) with matching
Senior ML Platform Engineer
Senior ML Platform Engineer

Echo • San Francisco (CA)

On-site
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1