Site Reliability Engineer (High Performance Computing)

Spacex

Hawthorne (CA)

On-site

USD 125,000 - 195,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options and equity potential
Comprehensive medical, vision, dental
401(k) retirement plan and paid leave

Job summary

SpaceX is seeking a Site Reliability Engineer for its HPC platform in Hawthorne, CA. You will own the end-to-end reliability of Linux-based infrastructure, IaC pipelines, storage, and user-facing services, collaborating with cross-functional teams to deliver scalable, repeatable operations.

Ideal candidates bring hands-on production infra experience, strong scripting, and a passion for reducing toil while ensuring mission-critical systems meet user needs and project deadlines.

Qualifications

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline or 2+ years of production infra experience.
  • 2+ years of Linux in production.
  • 2+ years operating production infrastructure including monitoring, debugging, and repair.

Responsibilities

  • Participate in on-call rotation and sustainable incident response.
  • Manage node lifecycle with infrastructure as code: OS images, firmware, configuration, kernel/driver stack.
  • Build observability for HPC admins and end users across cluster, node, and storage health.
  • Reduce toil with automation; split time between operating prod systems and writing tooling.
  • Sustainably manage resources including compute and storage.
  • Lead capacity planning with users to forecast tomorrow's compute/storage needs.
  • Collaborate with HPC systems engineers and other disciplines on maintainable infra.

Skills

Linux
Infrastructure as Code
Monitoring
Kubernetes
Python
DevOps
SRE

Education

Bachelor's degree in computer science, engineering, math, or a scientific discipline
2+ years of professional experience operating production infrastructure

Tools

Prometheus
Grafana
Nagios
Ansible
Puppet
Terraform
Docker
Kubernetes
Slurm/PBS/LSF

Job description

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal ofenableing human life on Mars.

SITE RELIABILITY ENGINEER (HIGH PERFORMANCE COMPUTING)

SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process.

We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications – the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts — you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just whether nodes ping. You’ll work alongside HPC systems engineers who design and commission clusters to help make them more reliable and provide world class services for world class engineers.

Aerospace experience is not required. We value engineers who treat teammates with fairness and respect, who are self-critical, and who will take ownership of hard production problems.

RESPONSIBILITIES:
  • Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortems
  • Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack
  • Build observability for both HPC administrators and end users — cluster, node, and storage health for operators, and job/workflow-level signal for the people running work on the platform
  • Reduce toil with automation; split time between operating production systems and writing the software that makes that work smaller
  • Sustainably manage resources, including compute and storage
  • Lead capacity planning with users across the company: understand what they will need next; and turn that into a concrete picture of tomorrow's compute and storage
  • Collaborate with HPC systems engineers and with engineers across all disciplines across the company on operable, maintainable infrastructure
BASIC QUALIFICATIONS:
  • Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree
  • 2+ years of experience with Linux operating systems in production
  • 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own
PREFERRED SKILLS AND EXPERIENCE:
  • 2+ years of professional experience in SRE, DevOps, or production infrastructure engineering
  • Experience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar)
  • Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar)
  • Experience writing scripts/code (eg. Python or similar languages) to automate common tasks
  • Experience with containers (Docker, Podman, Singularity/Apptainer)
  • Experience with Kubernetes administration for on-premise deployment
  • Experience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle management
  • Familiarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute — not required; we will teach this
  • Familiarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloads
  • Good understanding of version control, testing, continuous integration, build, deployment and monitoring
  • Ability to communicate clearly with users, peers, and vendors in both incident and design settings
  • Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities
  • Eligibility for access to classified material up to TS/SCI with polygraph
ADDITIONAL REQUIREMENTS:
  • Position is based in Hawthorne, CA and is primarily on-site
  • Must be able to participate in an on-call rotation
  • Must be willing to work extended hours and weekends as needed for incidents, cluster bring-up, and time-critical failures
COMPENSATION AND BENEFITS:

Pay Range:
Level 1: $125,000.00 - $160,000.00
Level 2: $145,000.00 - $195,000.00

Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience.

Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law.

ITAR REQUIREMENTS:
  • To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.

SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.

Applicants wishing to view a copy of SpaceX’s Aff…? EEOCompliance@spacex.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Site Reliability Engineer — HPC & Automation (Silicon Engineering)

SpaceX • Redmond (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Site Reliability Engineer - HPC & Automation (Silicon Engineering)
Site Reliability Engineer - HPC & Automation (Silicon Engineering)

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Sr. HPC Systems Engineer (High Performance Computing)
Sr. HPC Systems Engineer (High Performance Computing)

Spacex • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Medical, vision, dental
402(k) retirement plan
+1
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Stock options
Health, vision and dental coverage
401(k) retirement plan
+2
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

SpaceX • Brownsville (TX)

On-site
USD 140,000 - 200,000
Sr. High Performance Computing (HPC) Systems Engineer
Sr. High Performance Computing (HPC) Systems Engineer

Future Ventures • Town of Texas (WI)

On-site
USD 140,000 - 210,000
Production Engineer, Site Reliability (Application Software)
Production Engineer, Site Reliability (Application Software)

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Long-term cash awards
Health insurance
+6
Production Engineer, Site Reliability (Application Software)
Production Engineer, Site Reliability (Application Software)

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, and dental insurance
Site Reliability Engineer - Top Secret Clearance
Site Reliability Engineer - Top Secret Clearance

SpaceX • Hawthorne (CA)

On-site
USD 125,000 - 175,000
Comprehensive medical benefits
401(k) retirement plan
Paid parental leave
+1
Site Reliability Engineer (Application Software)
Site Reliability Engineer (Application Software)

SpaceX • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) retirement plan
Medical, vision, dental coverage
+1