Production Engineer, Compute Team Lead

Fluidstack

San Francisco (CA)

On-site

USD 269,000 - 335,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fluidstack is seeking a leader for the Data Center Operations Team to steward tens of thousands of GPUs and ensure high availability for customer workloads. The role emphasizes automation, end-to-end ownership, and fast, reliable operations at scale in a frontier compute environment.

Ideal candidates have led SRE or production engineering teams, shipped automated remediation, and can articulate how to improve availability metrics. Bonus: Kubernetes, Slurm, and HPC familiarity.

Qualifications

  • Led SRE or production engineering teams running large fleets.
  • Able to explain how to move an availability number.
  • Shipped automated remediation that retired a runbook.
  • Has hired and grown engineers; team growth track record.

Responsibilities

  • Lead the compute production engineering team keeping tens of thousands of GPUs serving customers.
  • Own fleet availability for compute: define the SLOs, build the tooling, and move the number.
  • Build automation for the node lifecycle, provisioning, health checks, remediation, return to service, without human touch.
  • Set the on-call and escalation model that keeps response sharp without burning the team out.

Skills

Led SRE teams
Production engineering
Automated remediation
Hiring and growing engineers
GPU/HPC familiarity

Tools

Kubernetes
Slurm
GPU HPC fleets

Job description

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it. We are singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.
  • Velocity. We drive everything forward as fast as possible.
  • First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
The Data Center Operations Team
Examples of key problems the team is working on
  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
Role Scope
  • Lead the compute production engineering team keeping tens of thousands of GPUs serving customers.
  • Own fleet availability for compute: define the SLOs, build the tooling, and move the number.
  • Build automation for the node lifecycle, provisioning, health checks, remediation, return to service, without human touch.
  • Set the on-call and escalation model that keeps response sharp without burning the team out.
What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've led SRE or production engineering teams running large fleets.
  • You've moved an availability number and can explain exactly how.
  • You've shipped automated remediation that retired a runbook.
  • You hire and grow strong engineers, and they say so.
  • Bonus: GPU or HPC fleets. Kubernetes or Slurm. Hardware failure analytics. Customer-facing reliability.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Compensation Range: $269K - $335K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Production Engineer, Compute Team Lead
Production Engineer, Compute Team Lead

Fluidstack • Seattle (WA), New York (NY), Austin (TX), San Francisco (CA)

On-site
USD 180,000 - 280,000
Production Engineer, IaaS Team Lead
Production Engineer, IaaS Team Lead

Fluidstack • Austin (TX)

On-site
USD 225,000 - 284,000
Production Engineer, IaaS Team Lead
Production Engineer, IaaS Team Lead

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Compute Engineer, Deployment
Compute Engineer, Deployment

Fluidstack • San Francisco (CA)

On-site
USD 150,000 - 250,000
Production Engineer, Facilities Team Lead
Production Engineer, Facilities Team Lead

Fluidstack • Austin (TX)

On-site
USD 208,000 - 269,000
Production Engineer, Compute (GPU)
Production Engineer, Compute (GPU)

Fluidstack • Austin (TX)

On-site
USD 175,000 - 300,000
Competitive total compensation
Retirement or pension plan
Health, dental, and vision insurance
+1
Production Engineer, Facilities Team Lead
Production Engineer, Facilities Team Lead

Fluidstack • Seattle (WA)

On-site
USD 225,000 - 284,000
Principal Operations Engineer, Reliability
Principal Operations Engineer, Reliability

Fluidstack • Austin (TX)

On-site
USD 220,000 - 260,000