Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. is seeking an infrastructure research engineer to design and build core systems enabling scalable, efficient training of large models for deployment and research. You’ll ensure fast, reliable experimentation and training so research teams can focus on science, not bottlenecks.

This role blends deep systems and performance expertise with curiosity for ML at scale. You’ll own the training stack end to end, driving performance from GPU cycles to scientific progress.

Qualifications

  • Bachelor’s degree or equivalent in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Strong engineering skills with performant, maintainable code and debugging in complex codebases.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
  • Bias for action with initiative to work across stacks and teams to ship improvements.

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.
  • Establish standards for reliability, maintainability, and security.
  • Collaborate with researchers and engineers to build scalable infrastructure.
  • Publish and share learnings through internal docs, open-source libraries, or technical reports.

Skills

Distributed systems
Performance optimization
Team collaboration
Deep learning frameworks (PyTorch/JAX)

Education

Bachelor’s degree or equivalent in CS/EE/ML

Tools

PyTorch
JAX
Megatron-LM
DeepSpeed
XLA

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure research engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Thinking Machines fast and reliable to ensure our research teams can focus on science, not system bottlenecks.

This role is ideal for someone who blends deep systems and performance expertise with a curiosity for machine learning at scale. You’ll take ownership of the training stack end to end, ensuring every GPU cycle drives scientific progress.

What You’ll Do
  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.

  • Develop high-performance optimizations to maximize throughput and efficiency.

  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.

  • Establish standards for reliability, maintainability, and security, ensuring systems are robust under rapid iteration.

  • Collaborate with researchers and engineers to build scalable infrastructure.

  • Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Past experience working on distributed training for the world’s largest models to make them stable, reliable, and performant.

  • Track record of improving research productivity through infrastructure design or process improvements.

  • Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.

Logistics
  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
ML Infrastructure Engineer
ML Infrastructure Engineer

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Systems Generalist
Software Engineer, Systems Generalist

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Visa sponsorship
Relocation support
Health benefits
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Research Engineer, Infrastructure, Kernels
Research Engineer, Infrastructure, Kernels

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Software Engineer, Data Infra
Software Engineer, Data Infra

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Post-Training
Research, Post-Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Research Tools
Software Engineer, Research Tools

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Site Reliability Engineer, Post Training
Site Reliability Engineer, Post Training

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Visa sponsorship
Relocation support
Unlimited PTO
+1
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1