GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. is seeking an engineer to design, build, and operate the GPU supercomputing environment that powers large-scale training and inference. You will deliver high-performant, reliable, and cost-efficient compute so researchers can move fast at scale.

This evergreen role is open on an ongoing basis to express interest. You may apply even if your experience aligns with future opportunities; we review applications continuously and reach out as roles open.

Qualifications

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large-scale clusters and container orchestration systems (Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.

Responsibilities

  • Operate and automate large GPU clusters including provisioning, imaging, and capacity planning.
  • Write software that abstracts cluster management and presents a unified interface for training and inference.
  • Extend scheduling/orchestration for topology-aware placement, preemption, quotas, and fair-share multi-tenancy.
  • Monitor and improve operational metrics of speed, reliability, and error recovery.
  • Build reliable storage and artifact paths for datasets, checkpoints, and logs with clear retention and lineage.
  • Partner with researchers to unblock scale runs and advise on parallelism and performance trade-offs.

Skills

Python
Rust
Kubernetes
Slurm

Education

Bachelor’s degree or equivalent experience in CS/Engineering

Tools

CUDA/NCCL
Performance profiling
Linux

Job description

Thinking Machines Lab Inc. is seeking an engineer to design, build, and operate the GPU supercomputing environment that powers large-scale training and inference. You will deliver high-performant, reliable, and cost-efficient compute so researchers can move fast at scale.

This evergreen role is open on an ongoing basis to express interest. You may apply even if your experience aligns with future opportunities; we review applications continuously and reach out as roles open.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Supercomputing
Software Engineer, Supercomputing

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
GPU Networking Engineer for Large-Scale AI Fabric
GPU Networking Engineer for Large-Scale AI Fabric

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Infra Engineer - Scalable Cloud Platform
Senior AI Infra Engineer - Scalable Cloud Platform

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Training Infra Engineer — GPU Clusters
Senior AI Training Infra Engineer — GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 300,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays