Staff GPU Compute & Infra Automation Engineer (Equity)

ATOMS Careers page

San Francisco (CA)

On-site

USD 224,000 - 284,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
Equity
Sick Time and Unlimited Flexible Time Off
Team lunch every Tuesday and Thursday

Job summary

Atoms is seeking a skilled professional to manage GPU training clusters in their San Francisco office. The ideal candidate will have over 6 years of experience with Kubernetes and GPU operations, alongside strong programming in Python. The role demands a focus on automation and reliability while working onsite five days a week.

Compensation ranges from $224,000 to $284,000 per year, with additional benefits including health insurance, 401(k), and equity options.

Qualifications

  • 6+ years of experience operating GPU compute on Kubernetes, with judgment to scale as demand grows.
  • Strong programming and scripting skills in Python, Go, or similar.
  • Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.
  • Comfort with bare-metal Linux environments, GPU hardware, and networking.
  • A bias toward automation, reliability, and operating critical systems well.

Responsibilities

  • Manage and automate GPU training clusters, including provisioning and lifecycle management.
  • Automate bare-metal bring-up to quickly and reliably add new machines.
  • Build software abstractions for training and simulation workloads.
  • Work at the hardware/software boundary to ensure speed and reliability.
  • Diagnose and resolve issues quickly during day-to-day operations.
  • Design infrastructure to scale from a smaller cluster to a larger fleet.

Skills

GPU compute operations
Python
Kubernetes
Infrastructure-as-Code tools
Linux environments

Job description

Atoms is seeking a skilled professional to manage GPU training clusters in their San Francisco office. The ideal candidate will have over 6 years of experience with Kubernetes and GPU operations, alongside strong programming in Python. The role demands a focus on automation and reliability while working onsite five days a week.

Compensation ranges from $224,000 to $284,000 per year, with additional benefits including health insurance, 401(k), and equity options.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GPU Cluster Automation Engineer
Staff GPU Cluster Automation Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Staff HPC Network Engineer - Onsite SF, Equity & PTO
Staff HPC Network Engineer - Onsite SF, Equity & PTO

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Staff Cluster Infrastructure Engineer
Staff Cluster Infrastructure Engineer

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+3
Staff Cluster Infrastructure Engineer
Staff Cluster Infrastructure Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Senior ML Infrastructure Engineer - Scalable GPU Training
Senior ML Infrastructure Engineer - Scalable GPU Training

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
Senior HPC Network Architect for GPU Compute
Senior HPC Network Architect for GPU Compute

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, Life Insurance
401(k)
Unlimited Flexible Time Off
+1
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6
Staff HPC Network Engineer
Staff HPC Network Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, Life Insurance
401(k)
Unlimited Flexible Time Off
+1
Staff HPC Network Engineer
Staff HPC Network Engineer

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Atoms • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+2