GPU Supercomputing Reliability Engineer — Unlimited PTO

Mosaic.tech

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines is hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system. You will diagnose hardware anomalies, track root causes to the hardware, and coordinate fixes with vendors so researchers can run at scale.

Based in San Francisco, this full-time role requires owning drivers, kernel surfaces, and diagnostics, plus automating fleet monitoring and reliability improvements across multi-disciplinary

Qualifications

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in Python or Rust and ability to own projects end-to-end.
  • Experience operating large-scale clusters and container orchestration (Kubernetes/Slurm).
  • Comfortable owning end-to-end projects and working across different stacks and teams.
  • Fluency in debugging hardware issues and collaborating with hardware vendors.

Responsibilities

  • Investigate, reproduce, and remediate issues across large GPU clusters.
  • Own the drivers, kernel surface, and diagnostics that span hardware, firmware, and OS.
  • Automate the monitoring of fleet reliability and analyze error rates to validate fixes.
  • Drive the firmware lifecycle: tracking, qualification, staged rollout, and regression analysis.
  • Engage vendors directly to get real fixes; manage RMA flows when hardware needs replacement.
  • Monitor and improve GPU hardware health signals and turn them into actionable improvements.
  • Write postmortems and vendor cases to move issues forward.

Skills

Python
Rust
Kubernetes
Slurm
Linux systems
Hardware debugging

Education

Bachelor’s degree or equivalent experience

Tools

Kubernetes
Slurm
BMC / iDRAC / IPMI / Redfish

Job description

Thinking Machines is hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system. You will diagnose hardware anomalies, track root causes to the hardware, and coordinate fixes with vendors so researchers can run at scale.

Based in San Francisco, this full-time role requires owning drivers, kernel surfaces, and diagnostics, plus automating fleet monitoring and reliability improvements across multi-disciplinary

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Engineer, Supercomputing
Reliability Engineer, Supercomputing

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
GPU Compute Engineer — Fleet Reliability & Automation
GPU Compute Engineer — Fleet Reliability & Automation

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
Software Engineer, Supercomputing
Software Engineer, Supercomputing

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Production Engineer, Compute - Automate & Scale GPU Fleet
Production Engineer, Compute - Automate & Scale GPU Fleet

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Competitive total compensation
Retirement or pension plan
Health, dental, and vision insurance
+1
Remote-Ready Compute Engineer for Scalable GPU Fleets
Remote-Ready Compute Engineer for Scalable GPU Fleets

Insight Global • Town of Texas (WI)

On-site
USD 120,000 - 180,000
GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals