Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. is seeking an infrastructure research engineer to design and build core systems enabling scalable, efficient training of large models for deployment and research. You’ll ensure fast, reliable experimentation and training so research teams can focus on science, not bottlenecks.

This role blends deep systems and performance expertise with curiosity for ML at scale. You’ll own the training stack end to end, driving performance from GPU cycles to scientific progress.

Qualifications

  • Bachelor’s degree or equivalent in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Strong engineering skills with performant, maintainable code and debugging in complex codebases.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
  • Bias for action with initiative to work across stacks and teams to ship improvements.

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.
  • Establish standards for reliability, maintainability, and security.
  • Collaborate with researchers and engineers to build scalable infrastructure.
  • Publish and share learnings through internal docs, open-source libraries, or technical reports.

Skills

Distributed systems
Performance optimization
Team collaboration
Deep learning frameworks (PyTorch/JAX)

Education

Bachelor’s degree or equivalent in CS/EE/ML

Tools

PyTorch
JAX
Megatron-LM
DeepSpeed
XLA

Job description

Thinking Machines Lab Inc. is seeking an infrastructure research engineer to design and build core systems enabling scalable, efficient training of large models for deployment and research. You’ll ensure fast, reliable experimentation and training so research teams can focus on science, not bottlenecks.

This role blends deep systems and performance expertise with curiosity for ML at scale. You’ll own the training stack end to end, driving performance from GPU cycles to scientific progress.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
ML Infra Engineer: Scale Distributed Training & Research
ML Infra Engineer: Scale Distributed Training & Research

Doist • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
ML Infra Engineer: Scale & Optimize Large-Scale Training
ML Infra Engineer: Scale & Optimize Large-Scale Training

Physical Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Platform Engineer — Infra for Research on GPU Fleets
ML Platform Engineer — Infra for Research on GPU Fleets

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 180,000
Senior Distributed ML Training Engineer
Senior Distributed ML Training Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 220,000 - 310,000
Stock options
Health/dental/vision insurance
Meals provided in office
+1
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000