Distributed Training Infra Engineer for Large Models

Kindredventures

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress.

Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling.

Qualifications

  • Experience with distributed training frameworks and techniques to train large foundation models.
  • Strong grasp of parallelism, memory optimization, mixed precision, and communication overlap.
  • Ability to profile and debug performance in complex codebases from framework internals down to kernels and collectives.
  • Deep understanding of PyTorch and JAX and their underlying system architectures.
  • Bonus: contributions to open-source ML infrastructure (e.g. PyTorch, Megatron-LM, DeepSpeed, XLA).

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs
  • Research and test parallelization strategies and numerical precision trade-offs across model scales
  • Analyze, profile, and debug low-level GPU operations to maximize throughput and hardware utilization
  • Build reusable frameworks for checkpointing, fault tolerance, and reproducibility that stay robust under rapid research iteration
  • Collaborate with researchers to bring novel model architectures from prototype to full scale
  • Stay up-to-date on research to bring new ideas to work

Skills

Distributed training
FSDP
DeepSpeed
Megatron
PyTorch
JAX/XLA
Performance profiling
Memory optimization
Mixed precision
Kernels/Collectives
Open-source contributions

Tools

N/A

Job description

Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress.

Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Distributed AI Training Architect
Distributed AI Training Architect

cerence • United States

On-site
USD 180,000 - 250,000
Distributed Training Engineer for Multimodal Models
Distributed Training Engineer for Multimodal Models

Luma • Redwood City (CA)

On-site
USD 210,000 - 260,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote Deep Learning Engineer - Distributed Training Scale
Remote Deep Learning Engineer - Distributed Training Scale

BairesDev • Peru (IL)

On-site
USD 120,000 - 180,000
Remote work
Competitive USD compensation
Home setup provided
+3
Infrastructure Research Engineer - Distributed AI Training
Infrastructure Research Engineer - Distributed AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000
Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1