GPU Systems Engineer for AI Training Clusters

Thinkingmachines

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines in San Francisco is looking for an engineer to design, build, and operate a GPU supercomputing environment. This role involves delivering reliable and efficient compute capabilities for extensive research and training applications.

The ideal candidate will possess a Bachelor’s degree and demonstrate proficiency in a backend language such as Python or Rust, with experience in large-scale clusters and container orchestration systems. Competitive benefits include generous health plans and unlimited PTO.

Qualifications

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems.

Responsibilities

  • Operate and automate large GPU clusters including provisioning, imaging, and capacity planning.
  • Write software that abstracts cluster management and presents a unified interface for training and inference.
  • Monitor and improve operational metrics of speed, reliability, and error recovery.

Skills

Backend language proficiency (Python or Rust)
Cluster operation experience
Container orchestration systems knowledge (Kubernetes or Slurm)
Systems background (Linux, networking)

Education

Bachelor’s degree or equivalent experience

Job description

Thinking Machines in San Francisco is looking for an engineer to design, build, and operate a GPU supercomputing environment. This role involves delivering reliable and efficient compute capabilities for extensive research and training applications.

The ideal candidate will possess a Bachelor’s degree and demonstrate proficiency in a backend language such as Python or Rust, with experience in large-scale clusters and container orchestration systems. Competitive benefits include generous health plans and unlimited PTO.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Reliability Engineer - AI Supercomputing Fleet
GPU Reliability Engineer - AI Supercomputing Fleet

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Paid parental leave
Relocation support
+1
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Software Engineer, Supercomputing
Software Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior AI GPU Cluster Architect
Senior AI GPU Cluster Architect

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Infra Engineer – Large-Scale GPU Training
Staff Infra Engineer – Large-Scale GPU Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays