GPU Reliability Engineer - AI Supercomputing Fleet

Thinkingmachines

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Paid parental leave
Relocation support
Health, dental, and vision benefits

Job summary

Thinking Machines in San Francisco is hiring an engineer to ensure the reliability of their GPU supercomputing fleet. You'll be responsible for diagnosing hardware issues and collaborating with vendors to resolve them efficiently.

Ideal candidates hold a Bachelor’s degree in computer science or engineering, possess backend programming skills in Python or Rust, and have experience with large-scale systems. The position offers a competitive salary range of $350,000 to $475,000 and generous benefits including health insurance and unlimited PTO.

Qualifications

  • Experience diagnosing and resolving hardware issues with GPU clusters.
  • Ability to automate the monitoring of fleet reliability.
  • Track and improve GPU hardware health signals.

Responsibilities

  • Investigate and remediate issues across large GPU clusters.
  • Own the drivers, kernel surface, and diagnostics spanning hardware and software.
  • Engage effectively with vendors for issue resolution.

Skills

Proficiency in at least one backend language (Python or Rust)
Experience operating large-scale clusters and container orchestration systems (Kubernetes or Slurm)
Linux systems and debugging tools
Strong writing skills for vendor cases
Statistical rigor in analyzing reliability
Experience engaging hardware vendors directly

Education

Bachelor’s degree or equivalent experience in computer science, engineering, or similar

Job description

Thinking Machines in San Francisco is hiring an engineer to ensure the reliability of their GPU supercomputing fleet. You'll be responsible for diagnosing hardware issues and collaborating with vendors to resolve them efficiently.

Ideal candidates hold a Bachelor’s degree in computer science or engineering, possess backend programming skills in Python or Rust, and have experience with large-scale systems. The position offers a competitive salary range of $350,000 to $475,000 and generous benefits including health insurance and unlimited PTO.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer, Supercomputing
Reliability Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Paid parental leave
Relocation support
+1
GPU Systems Engineer for AI Training Clusters
GPU Systems Engineer for AI Training Clusters

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Software Engineer, Supercomputing
Software Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Senior GPU HPC Platform Reliability Engineer
Senior GPU HPC Platform Reliability Engineer

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Senior GPU Fleet Reliability Engineer
Senior GPU Fleet Reliability Engineer

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Learning and growth opportunities
Visa sponsorship and relocation assistance
+1
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Senior Software Engineer, GPU Fleet Reliability
Senior Software Engineer, GPU Fleet Reliability

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 300,000
Health benefits
Paid time off
401(k) match
+1