ML Infra Tech Lead: Scalable Training & Inference

Reducto

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Daily Lunch
Commuter Reimbursement
Comprehensive Insurance
Health and Wellness Budget
Parental Leave

Job summary

Reducto in San Francisco is seeking an ML Infrastructure Tech Lead to own and scale high-performance model training and inference systems. This hands-on role emphasizes building, debugging, and optimizing the production stack while guiding architectural decisions.

You'll work across kernels, runtimes, batching, and distributed training with Kubernetes, collaborating closely with ML and Platform teams. This in-person role requires delivering impact quickly in a fast-moving startup.

Qualifications

  • 5+ years building production infrastructure with ML systems.
  • Led complex technical projects from ambiguous problems through production.
  • Able to set direction and implement the hardest parts.
  • Strong Python and systems-engineering skills.
  • Understanding of GPU training/inference performance.
  • Comfortable with Kubernetes and distributed frameworks.
  • Ability to reason across low-level performance and high-level architecture.
  • High bar for quality, precision, and reliability.
  • Thrives in fast-changing, high-growth environments.
  • Takes ownership from strategy to execution.

Responsibilities

  • Own the technical direction and roadmap for Reducto's ML infrastructure.
  • Build and maintain training and inference stack balancing speed and production.
  • Optimize model serving at kernels, runtimes, batching, scheduling, and distribution.
  • Design systems for multi-node, multi-GPU training and inference.
  • Improve GPU utilization, latency, throughput, reliability, observability, cost.
  • Develop benchmarks to identify bottlenecks and guide investments.
  • Evaluate advances in training and inference and apply the ones that matter.
  • Build tooling and abstractions for moving from experiments to production.
  • Partner with ML and Platform teams on architecture and prioritization.
  • Raise the engineering bar through design reviews, mentorship, and leadership.

Skills

Python
Systems engineering
GPU training
Kubernetes
Distributed training
Production infra
Leadership
Ownership

Tools

CUDA
Triton
PyTorch
TensorRT-LLM
vLLM
Ray

Job description

Reducto in San Francisco is seeking an ML Infrastructure Tech Lead to own and scale high-performance model training and inference systems. This hands-on role emphasizes building, debugging, and optimizing the production stack while guiding architectural decisions.

You'll work across kernels, runtimes, batching, and distributed training with Kubernetes, collaborating closely with ML and Platform teams. This in-person role requires delivering impact quickly in a fast-moving startup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
ML Infra Engineer: Build Scalable ML Services & MLOps
ML Infra Engineer: Build Scalable ML Services & MLOps

Stripe • United States

Hybrid
CAD 172,000 - 258,000
Equity
Retirement plans
Health benefits
+1
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Infrastructure Engineer — AI/ML Reliability & Scale
Infrastructure Engineer — AI/ML Reliability & Scale

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free daily lunch
Reimbursed transportation
+3
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO
Machine Learning Infra Engineer
Machine Learning Infra Engineer

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
Cloud-Scale Backend Engineer for ML Inference
Cloud-Scale Backend Engineer for ML Inference

Praxis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 250,000
Senior Lead, Scalable ML Inference Infrastructure
Senior Lead, Scalable ML Inference Infrastructure

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3