Senior SRE: AI/GPU Infra & On-Prem Clusters

SpaceX

Palo Alto (CA)

On-site

USD 165,000 - 265,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
Long-term incentives
Medical/Vision/Dental
401(k)
Parental leave
Paid vacation
Paid holidays

Job summary

SpaceX is seeking a Sr. Site Reliability Engineer for Starshield, working on accelerator-grade GPU/CPU infrastructure and on-prem AI clusters.

You will design, operate, and scale critical systems across Top Secret data centers while collaborating with AI engineers to ensure highly available services. The role requires expert Linux, Kubernetes, and infrastructure tool experience, plus the ability to mentor junior engineers and drive technical excellence in a fast-paced defense‑related program.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or engineering, or 7+ years in lieu.
  • 5+ years of experience with Linux operating systems.
  • 5+ years of experience with Kubernetes.
  • Experience with Terraform, Ansible, or other infrastructure tools.
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes).
  • Experience scripting in Bash, Python, or other similar languages; development in Python, C++, or Go.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU as a service for external customers on bare metal hardware and virtualized platforms.
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale).
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems.
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage.
  • Collaborate with AI engineers to create scalable, operable products.
  • Mentor and train junior engineers; lead team to technical excellence.
  • Engage in and improve the whole lifecycle of services from inception to operation.
  • Monitor and alert supporting systems to ensure high availability.

Skills

Linux administration
Kubernetes
Python
Go
Bash scripting

Education

Bachelor’s degree in CS/IT or engineering

Tools

Terraform
Ansible
OCI containers / Docker

Job description

SpaceX is seeking a Sr. Site Reliability Engineer for Starshield, working on accelerator-grade GPU/CPU infrastructure and on-prem AI clusters.

You will design, operate, and scale critical systems across Top Secret data centers while collaborating with AI engineers to ensure highly available services. The role requires expert Linux, Kubernetes, and infrastructure tool experience, plus the ability to mentor junior engineers and drive technical excellence in a fast-paced defense‑related program.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
401(k) plan
Medical, vision, dental coverage
+2
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Senior SRE for AI GPU Infra (Top Secret)
Senior SRE for AI GPU Infra (Top Secret)

SpaceX • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options and long-term incentives
Comprehensive medical/dental/vision
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
+4