Site Reliability Engineer - HPC Compute Platforms

Spacex

Hawthorne (CA)

On-site

USD 125,000 - 195,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock options and equity potential
Comprehensive medical, vision, dental
401(k) retirement plan and paid leave

Job summary

SpaceX is seeking a Site Reliability Engineer for its HPC platform in Hawthorne, CA. You will own the end-to-end reliability of Linux-based infrastructure, IaC pipelines, storage, and user-facing services, collaborating with cross-functional teams to deliver scalable, repeatable operations.

Ideal candidates bring hands-on production infra experience, strong scripting, and a passion for reducing toil while ensuring mission-critical systems meet user needs and project deadlines.

Qualifications

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline or 2+ years of production infra experience.
  • 2+ years of Linux in production.
  • 2+ years operating production infrastructure including monitoring, debugging, and repair.

Responsibilities

  • Participate in on-call rotation and sustainable incident response.
  • Manage node lifecycle with infrastructure as code: OS images, firmware, configuration, kernel/driver stack.
  • Build observability for HPC admins and end users across cluster, node, and storage health.
  • Reduce toil with automation; split time between operating prod systems and writing tooling.
  • Sustainably manage resources including compute and storage.
  • Lead capacity planning with users to forecast tomorrow's compute/storage needs.
  • Collaborate with HPC systems engineers and other disciplines on maintainable infra.

Skills

Linux
Infrastructure as Code
Monitoring
Kubernetes
Python
DevOps
SRE

Education

Bachelor's degree in computer science, engineering, math, or a scientific discipline
2+ years of professional experience operating production infrastructure

Tools

Prometheus
Grafana
Nagios
Ansible
Puppet
Terraform
Docker
Kubernetes
Slurm/PBS/LSF

Job description

SpaceX is seeking a Site Reliability Engineer for its HPC platform in Hawthorne, CA. You will own the end-to-end reliability of Linux-based infrastructure, IaC pipelines, storage, and user-facing services, collaborating with cross-functional teams to deliver scalable, repeatable operations.

Ideal candidates bring hands-on production infra experience, strong scripting, and a passion for reducing toil while ensuring mission-critical systems meet user needs and project deadlines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE — HPC Platform Engineer (Compute & ML)
SRE — HPC Platform Engineer (Compute & ML)

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options and long-term incentives
Bonuses and discretionary incentives
401(k) retirement plan
+5
Site Reliability Engineer: HPC & Automation
Site Reliability Engineer: HPC & Automation

SpaceX • Redmond (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Site Reliability Engineer (High Performance Computing)
Site Reliability Engineer (High Performance Computing)

Spacex • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options and equity potential
Comprehensive medical, vision, dental
401(k) retirement plan and paid leave
Site Reliability Engineer – HPC & Rocket Engine Systems
Site Reliability Engineer – HPC & Rocket Engine Systems

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 175,000
Stock options
Performance bonuses
Medical, vision, dental coverage
+1
Production Site Reliability Engineer—Mission-Critical Infra
Production Site Reliability Engineer—Mission-Critical Infra

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, and dental insurance
Site Reliability Engineer (High Performance Computing)
Site Reliability Engineer (High Performance Computing)

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options and long-term incentives
Bonuses and discretionary incentives
401(k) retirement plan
+5
HPC Site Reliability Engineer — Automation & Infra
HPC Site Reliability Engineer — Automation & Infra

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA)

On-site
USD 125,000 - 150,000
Health insurance
401(k) retirement plan
Paid vacation and holidays
+1
Senior Site Reliability Engineer - Cloud & Automation
Senior Site Reliability Engineer - Cloud & Automation

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+1
SRE for HPC & Automation - Scale Silicon Workflows
SRE for HPC & Automation - Scale Silicon Workflows

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Site Reliability Engineer — Mission-Critical Systems
Site Reliability Engineer — Mission-Critical Systems

Iac/interactivecorp • Sacramento (CA)

On-site
USD 120,000 - 170,000
Stock options
Health insurance
Paid vacation