Staff Cluster Infrastructure Engineer

ATOMS Careers page

San Francisco (CA)

On-site

USD 224,000 - 284,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
Equity
Sick Time and Unlimited Flexible Time Off
Team lunch every Tuesday and Thursday

Job summary

Atoms is seeking a skilled professional to manage GPU training clusters in their San Francisco office. The ideal candidate will have over 6 years of experience with Kubernetes and GPU operations, alongside strong programming in Python. The role demands a focus on automation and reliability while working onsite five days a week.

Compensation ranges from $224,000 to $284,000 per year, with additional benefits including health insurance, 401(k), and equity options.

Qualifications

  • 6+ years of experience operating GPU compute on Kubernetes, with judgment to scale as demand grows.
  • Strong programming and scripting skills in Python, Go, or similar.
  • Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.
  • Comfort with bare-metal Linux environments, GPU hardware, and networking.
  • A bias toward automation, reliability, and operating critical systems well.

Responsibilities

  • Manage and automate GPU training clusters, including provisioning and lifecycle management.
  • Automate bare-metal bring-up to quickly and reliably add new machines.
  • Build software abstractions for training and simulation workloads.
  • Work at the hardware/software boundary to ensure speed and reliability.
  • Diagnose and resolve issues quickly during day-to-day operations.
  • Design infrastructure to scale from a smaller cluster to a larger fleet.

Skills

GPU compute operations
Python
Kubernetes
Infrastructure-as-Code tools
Linux environments

Job description

Atoms is building the machines that power the next era of progress. Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far less intelligent, far less efficient, and far more constrained. We’re changing that.

Atoms builds Physical AI - real-world robots for the industries that move civilization forward, starting with food, mining, and transport. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.

This work requires more than robotics. It requires deep integration across hardware, software, AI, operations, manufacturing, and real estate. We don’t just build machines in a lab. We deploy them into real environments, operate them, learn from them, and improve them until they work at scale.

We are roboticists, engineers, operators, and builders. We believe the next great technology companies will not only transform information, but the physical systems that shape everyday life.

What you’ll do
  • Manage and automate our GPU training clusters, including provisioning, bootstrapping, and lifecycle management.
  • Automate bare‑metal bring‑up so new machines come online quickly and reliably as we add capacity.
  • Build software abstractions that present a clean, unified interface to our training and simulation workloads.
  • Work at the hardware/software boundary, where speed and reliability are critical, continuously raising the bar for automation and uptime.
  • Run day‑to‑day operations: diagnose and resolve issues quickly when systems are under pressure.
  • Design our infrastructure to scale smoothly as we grow from a smaller cluster of machines toward a larger fleet.
What we’re looking for
  • 6+ years experience operating GPU compute on Kubernetes (or similar orchestration), with the judgment to scale it as demand grows.
  • Strong programming and scripting skills in Python, Go, or similar.
  • Familiarity with Infrastructure‑as‑Code tools such as Terraform or CloudFormation.
  • Comfort with bare‑metal Linux environments, GPU hardware, and networking.
  • A bias toward automation, reliability, and operating critical systems well.

At Atoms, you’ll work on one of the defining challenges of our time — bringing automation into the physical world to drive real, lasting impact. We exist to uncover valuable unknown truths and turn them into progress, which means constantly pushing beyond what’s known and building what doesn’t yet exist. The work is ambitious and often challenging, but it’s grounded in a shared sense of purpose and a team committed to seeing it through together. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team — so we invest in both, creating an environment where you can do your best work and grow.

What else you need to know

This role is based in our San Francisco office. Atoms is a company driven by invention and continuous change – we are constantly reimagining our industries, building new products, and refining how we operate. We do our best work together. That’s why all of our office‑based teams work onsite, five days a week.

Base salary range for this role: $224,000 - $284,000 per year. Actual compensation will be determined on an individual basis and may vary depending on experience, skills, and qualifications.

Base salary is just one part of your total rewards package. You may also be eligible for equity awards and an annual performance‑based bonus.

Benefits Summary (USA Full‑Time Exempt Employees)
  • Medical, Dental, Vision, Disability, and Life Insurance
  • Flexible Spending Account / Health Savings Account Options
  • 401(k)
  • Equity
  • Sick Time, Unlimited Flexible Time Off, and Paid Holidays
  • Team lunch in our SoMa office every Tuesday and Thursday

Benefits are subject to change at the company's discretion. Atoms accepts applications on an ongoing basis.

As set forth in ATOMS Careers page’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Cluster Infrastructure Engineer
Staff Cluster Infrastructure Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Staff Cluster Infrastructure Engineer
Staff Cluster Infrastructure Engineer

Cssmerge • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, and Life Insurance
401(k)
Unlimited Flexible Time Off
Staff HPC Network Engineer
Staff HPC Network Engineer

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Staff HPC Network Engineer
Staff HPC Network Engineer

Cssmerge • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, and Life Insurance
401(k)
Unlimited Flexible Time Off and Paid Holidays
+1
Staff HPC Network Engineer
Staff HPC Network Engineer

Atoms • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, Life Insurance
401(k)
Unlimited Flexible Time Off
+1
Infrastructure Product Manager
Infrastructure Product Manager

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical insurance
Dental insurance
Vision insurance
+4
Staff Backend Engineer
Staff Backend Engineer

jobr.pro • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+3
Staff Backend Engineer
Staff Backend Engineer

ATOMS Careers page • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision, Disability, and Life Insurance
Flexible Spending Account / Health Savings Account Options
401(k)
+3
Staff Backend Engineer
Staff Backend Engineer

Cssmerge • San Francisco (CA)

On-site
USD 224,000 - 284,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off
+1
Senior Full Stack Engineer
Senior Full Stack Engineer

ATOMS Careers page • San Francisco (CA)

On-site
USD 176,000 - 242,000
Medical, Dental, Vision Insurance
401(k)
Unlimited Flexible Time Off