ML Infra Tech Lead: High-Performance GPU & K8s

Reducto

Santa Fe (NM)

On-site

USD 190,000 - 230,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Daily Lunch
Commuter Reimbursement
Comprehensive Insurance
Health and Wellness Budget
Parental Leave

Job summary

Reducto seeks an ML Infrastructure Tech Lead to own and optimize our high-performance infrastructure for model training and inference. You will drive architecture choices, implement scalable multi-GPU systems, and partner with ML and Platform teams to deliver production-ready software.

This hands-on leadership role requires deep Python and systems engineering skills, extensive experience with GPUs, and the ability to balance rapid experimentation with reliable production serving.

Qualifications

  • Have 5+ years of experience building production infrastructure, including significant ML systems experience.
  • Have led complex technical projects from an ambiguous problem through production deployment.
  • Are equally comfortable setting direction and personally implementing the hardest parts.
  • Have strong Python and systems-engineering skills.
  • Understand the performance characteristics of modern GPU training or inference workloads.
  • Are comfortable with Kubernetes and distributed training or serving frameworks.
  • Can reason across low-level model performance and higher-level platform architecture.
  • Hold yourself to a high bar for quality, precision, and operational reliability.
  • Operate well in a fast-changing, high-growth environment.
  • Take full ownership from strategy through execution.

Responsibilities

  • Own the technical direction and roadmap for Reducto's ML infrastructure.
  • Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving.
  • Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference.
  • Design systems for reliable multi-node, multi-GPU training and inference.
  • Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
  • Develop benchmarks that identify bottlenecks and guide infrastructure investments.
  • Evaluate state-of-the-art advances in training and inference and apply the ones that matter.
  • Build the tooling and abstractions that help ML engineers move quickly from experiments to production.
  • Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization.
  • Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership.

Skills

Python
Systems engineering
Kubernetes
Distributed training
GPU workloads
ML infrastructure

Tools

CUDA
Triton
PyTorch

Job description

Reducto seeks an ML Infrastructure Tech Lead to own and optimize our high-performance infrastructure for model training and inference. You will drive architecture choices, implement scalable multi-GPU systems, and partner with ML and Platform teams to deliver production-ready software.

This hands-on leadership role requires deep Python and systems engineering skills, extensive experience with GPUs, and the ability to balance rapid experimentation with reliable production serving.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Lead ML Infrastructure Engineer (Kubernetes + GPUs)
Lead ML Infrastructure Engineer (Kubernetes + GPUs)

Cohere • California (MO)

Hybrid
USD 180,000 - 260,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • Santa Fe (NM)

On-site
USD 190,000 - 230,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Senior ML Infra Platform Engineer — Kubernetes & GPUs
Senior ML Infra Platform Engineer — Kubernetes & GPUs

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Infra Engineer: GPU Fleet & Inference Orchestrator
ML Infra Engineer: GPU Fleet & Inference Orchestrator

Generalist • San Francisco (CA)

On-site
USD 120,000 - 160,000
Founding ML Infra Architect
Founding ML Infra Architect

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2