Staff GPU Cluster Automation Engineer

Atoms

San Francisco (CA)

On-site

USD 224,000 - 284,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
Team lunch every Tuesday and Thursday

Job summary

Atoms is seeking an experienced engineer in San Francisco to manage GPU training clusters and automate operations. You will work with hardware and software at a critical level.

The ideal candidate has over 6 years of experience with GPU compute on Kubernetes and strong programming skills in Python or Go. This role offers a competitive salary and comprehensive benefits including medical, dental, and 401(k).

Qualifications

  • 6+ years experience operating GPU compute on Kubernetes.
  • Strong programming and scripting skills in Python, Go, or similar.
  • Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.

Responsibilities

  • Manage and automate GPU training clusters and lifecycle management.
  • Automate bare-metal bring-up quickly and reliably.
  • Build software abstractions for training and simulation workloads.

Skills

Operating GPU compute on Kubernetes
Programming and scripting in Python or Go
Infrastructure-as-Code tools
Bare-metal Linux environments
Automation and reliability

Job description

Atoms is seeking an experienced engineer in San Francisco to manage GPU training clusters and automate operations. You will work with hardware and software at a critical level.

The ideal candidate has over 6 years of experience with GPU compute on Kubernetes and strong programming skills in Python or Go. This role offers a competitive salary and comprehensive benefits including medical, dental, and 401(k).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Cluster Infrastructure Engineer
Staff Cluster Infrastructure Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Senior HPC Network Architect for GPU Compute
Senior HPC Network Architect for GPU Compute

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, Life Insurance
401(k)
Unlimited Flexible Time Off
+1
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2
Staff HPC Network Engineer
Staff HPC Network Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, Life Insurance
401(k)
Unlimited Flexible Time Off
+1
Lead HPC/GPU Cluster Architect | Automate & Scale
Lead HPC/GPU Cluster Architect | Automate & Scale

The San Francisco Compute Company • Boston (MA)

Hybrid
USD 140,000 - 200,000
Generous equity grant
Visa Sponsorships
Retirement matching
+5
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Production Engineer, Compute - Automate & Scale GPU Fleet
Production Engineer, Compute - Automate & Scale GPU Fleet

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Competitive total compensation
Retirement or pension plan
Health, dental, and vision insurance
+1
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Backend Platform Engineer – GPU Cluster Automation
Backend Platform Engineer – GPU Cluster Automation

TensorWave • Las Vegas (NV)

On-site
USD 140,000 - 210,000
Stock Options
Excellent health insurance
401(k)
+6