Senior SRE: AI Infrastructure & GPU Systems

SpaceX

Hawthorne (CA)

On-site

USD 165,000 - 265,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) plan
Paid time off
Health insurance

Job summary

SpaceX in Hawthorne, CA, is hiring a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale the GPU and software infrastructure for the US government Starshield constellation.

This senior role leads deployments to secure data centers, builds automation with Kubernetes, Terraform, and Ansible, and mentors engineers while ensuring high availability and strict security clearances.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline with 5+ years Linux/Kubernetes experience.
  • 5+ years of Kubernetes experience.
  • 5+ years of Linux OS experience.
  • Experience with infrastructure tools like Terraform, Ansible.
  • Experience with containerization (OCI/Kubernetes).
  • Scripting in Bash, Python, or similar languages.
  • Development experience in Python, C++, or Go.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support for external customers on bare metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions (100k+ GPU scale).
  • Develop automation to deploy and manage on-prem Kubernetes/AI clusters and OSes.
  • Deploy and manage core infra: databases, monitoring, storage.
  • Collaborate with AI engineers to build scalable, operable products.
  • Own the full lifecycle from design to deployment and refinement.
  • Ensure high availability via monitoring and alerts.
  • Mentor and train junior engineers; drive technical excellence.

Skills

Kubernetes
Linux
Terraform/Ansible
Containerization
Python/Bash
Go/C++

Education

Bachelor’s degree in CS/IT/Engineering

Tools

Terraform
Ansible
Kubernetes
OCI containers

Job description

SpaceX in Hawthorne, CA, is hiring a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale the GPU and software infrastructure for the US government Starshield constellation.

This senior role leads deployments to secure data centers, builds automation with Kubernetes, Terraform, and Ansible, and mentors engineers while ensuring high availability and strict security clearances.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE for AI GPU Infra (Top Secret)
Senior SRE for AI GPU Infra (Top Secret)

SpaceX • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options and long-term incentives
Comprehensive medical/dental/vision
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Site Reliability Engineer — AI Infra & GPU Clusters
Site Reliability Engineer — AI Infra & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k)
Health coverage
+1
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3