Distributed Training Infra Engineer for Large Models

Kindredventures

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress.

Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling.

Qualifications

  • Experience with distributed training frameworks and techniques to train large foundation models.
  • Strong grasp of parallelism, memory optimization, mixed precision, and communication overlap.
  • Ability to profile and debug performance in complex codebases from framework internals down to kernels and collectives.
  • Deep understanding of PyTorch and JAX and their underlying system architectures.
  • Bonus: contributions to open-source ML infrastructure (e.g. PyTorch, Megatron-LM, DeepSpeed, XLA).

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs
  • Research and test parallelization strategies and numerical precision trade-offs across model scales
  • Analyze, profile, and debug low-level GPU operations to maximize throughput and hardware utilization
  • Build reusable frameworks for checkpointing, fault tolerance, and reproducibility that stay robust under rapid research iteration
  • Collaborate with researchers to bring novel model architectures from prototype to full scale
  • Stay up-to-date on research to bring new ideas to work

Skills

Distributed training
FSDP
DeepSpeed
Megatron
PyTorch
JAX/XLA
Performance profiling
Memory optimization
Mixed precision
Kernels/Collectives
Open-source contributions

Tools

N/A

Job description

Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress.

Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: High-Performance Distributed Training
ML Systems Engineer: High-Performance Distributed Training

Motional • San Francisco (CA)

Hybrid
USD 144,000 - 192,000
Medical
Dental
Vision
+4
Member of Technical Staff: Training Infrastructure
Member of Technical Staff: Training Infrastructure

Wintermeyer Ventures • San Francisco (CA)

On-site
USD 200,000 - 375,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
Staff Training Infra Engineer for Scalable ML Pipelines
Staff Training Infra Engineer for Scalable ML Pipelines

WeHireYou • Paris (IN)

Hybrid
USD 102,000 - 147,000
Lunch stipend
Health & dental benefits
Parental leave
+4