Distributed Systems Engineer for AI Clusters

Thinking Machines Lab Inc.

San Francisco (CA)

On-site

USD 300,000 - 400,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines is hiring a Software Engineer, Distributed Systems in San Francisco to design and build core distributed systems for orchestration, scheduling, storage, and networking across thousands of machines. You will support Inkling's training clusters and Tinker's serving platform with fault-tolerant, scalable software.

You’ll work on consensus, fault tolerance, and performance under real-world failure conditions, collaborating with research and infrastructure teams to solve hard

Qualifications

  • 5+ years of experience building large-scale distributed systems.
  • Proficiency in Python and Go, C++, or another systems-level language.
  • Strong understanding of distributed systems fundamentals: consensus, consistency, replication, and fault tolerance.
  • Experience with network programming, load balancing, or distributed storage systems.

Responsibilities

  • Design and build distributed systems for compute orchestration, scheduling, storage, and networking across large GPU and TPU clusters.
  • Develop fault-tolerant systems that keep running correctly as hardware fails, networks partition, and workloads scale.
  • Build the distributed storage and data orchestration layers that move and persist large volumes of training and model data.
  • Improve the performance and efficiency of collective communication, scheduling, and resource allocation across thousands of machines.
  • Partner with research and infrastructure teams to identify systems bottlenecks and design solutions from first principles.
  • Write production-quality code and help shape the architecture of systems used company-wide.

Skills

Python
Go
C++
Distributed systems
Network programming
Fault tolerance

Job description

Thinking Machines is hiring a Software Engineer, Distributed Systems in San Francisco to design and build core distributed systems for orchestration, scheduling, storage, and networking across thousands of machines. You will support Inkling's training clusters and Tinker's serving platform with fault-tolerant, scalable software.

You’ll work on consensus, fault tolerance, and performance under real-world failure conditions, collaborating with research and infrastructure teams to solve hard

Get your free, confidential resume review.

or drag and drop your file here.