SRE — HPC Platform Engineer (Compute & ML)

InvestedintheMission

Hawthorne (CA)

On-site

USD 125,000 - 195,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options and long-term incentives
Bonuses and discretionary incentives
401(k) retirement plan
Medical, vision, dental coverage
Paid parental leave
3 weeks paid vacation
10+ paid holidays per year
Sick leave benefits

Job summary

SpaceX in Hawthorne, CA is seeking a Site Reliability Engineer for High Performance Computing to own the production infrastructure as a product. This role emphasizes reducing toil, automating workflows, and building observability for HPC clusters, storage, and user-facing apps.

You will work with HPC systems engineers to improve reliability, operate real infrastructure, and write code to automate common tasks. A background in Linux, IaC, and on-call incident response is essential.

Qualifications

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline, or 2+ years of infra experience in lieu of a degree.
  • 2+ years of experience with Linux operating systems in production.
  • 2+ years operating production infrastructure (servers, services, or networks), including monitoring and debugging.

Responsibilities

  • Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortems.
  • Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack.
  • Build observability for HPC admins and end users — cluster, node, and storage health.
  • Reduce toil with automation; split time between operating production systems and writing software to automate tasks.
  • Sustainably manage resources, including compute and storage.
  • Lead capacity planning with users across the company: forecast tomorrow's compute/storage needs.
  • Collaborate with HPC systems engineers and engineers across disciplines on maintainable infrastructure.

Skills

SRE/DevOps
Monitoring/Alerting
Infrastructure as Code
Python scripting
Containers
Kubernetes (on-prem)
Distributed storage
HPC familiarity
Scientific computing/ML
Version control/CI
Communication

Education

Bachelor's degree in CS/Engineering/Math or scientific discipline
2+ years infra experience in lieu of degree

Tools

Prometheus
Grafana
Nagios
Ansible
Puppet
Terraform
Docker
Podman
Singularity/Apptainer
Kubernetes
Slurm/PBS/LSF
PyTorch
TensorFlow
Git

Job description

SpaceX in Hawthorne, CA is seeking a Site Reliability Engineer for High Performance Computing to own the production infrastructure as a product. This role emphasizes reducing toil, automating workflows, and building observability for HPC clusters, storage, and user-facing apps.

You will work with HPC systems engineers to improve reliability, operate real infrastructure, and write code to automate common tasks. A background in Linux, IaC, and on-call incident response is essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - HPC Compute Platforms
Site Reliability Engineer - HPC Compute Platforms

Spacex • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options and equity potential
Comprehensive medical, vision, dental
401(k) retirement plan and paid leave
SRE for HPC & Automation - Scale Silicon Workflows
SRE for HPC & Automation - Scale Silicon Workflows

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
SRE: Mission-Critical Vehicle Software
SRE: Mission-Critical Vehicle Software

SpaceX • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Discretionary bonuses
Medical, vision, and dental coverage
+1
Production Site Reliability Engineer—Mission-Critical Infra
Production Site Reliability Engineer—Mission-Critical Infra

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, and dental insurance
Production SRE for Mission-Critical Vehicle Software
Production SRE for Mission-Critical Vehicle Software

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Long-term cash awards
Health insurance
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
Site Reliability Engineer: HPC & Automation
Site Reliability Engineer: HPC & Automation

SpaceX • Redmond (WA)

On-site
USD 125,000 - 150,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Senior Mission-Critical SRE for Vehicle Software
Senior Mission-Critical SRE for Vehicle Software

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2
Site Reliability Engineer (High Performance Computing)
Site Reliability Engineer (High Performance Computing)

Spacex • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options and equity potential
Comprehensive medical, vision, dental
401(k) retirement plan and paid leave
Senior SRE — Starshield Cloud & Reliability
Senior SRE — Starshield Cloud & Reliability

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+1