Senior SRE: AI Infra & GPU Clusters

InvestedintheMission

Redmond (WA)

On-site

USD 165,000 - 270,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) plan
Medical, vision, dental coverage
Paid time off
Company shuttle

Job summary

SpaceX is seeking a Sr. Site Reliability Engineer (Starshield) to design, operate, and scale GPU/CPU infrastructure for critical national security data centers in a Starshield environment.

You will automate on-prem Kubernetes/AI clusters, deploy databases and monitoring, and collaborate with AI engineers to deliver scalable, reliable software products. The role requires Top Secret/DOE-level clearance, 5+ years in Linux/Kubernetes, and strong scripting in Bash or Python.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline, or 7+ years of relevant experience.
  • 5+ years of professional Linux experience.
  • 5+ years of Kubernetes experience.
  • Experience with Terraform, Ansible, or similar infrastructure tools.
  • Experience with containerization (OCI containers, Kubernetes).
  • Scripting in Bash or Python.
  • Development in Python, C++, or Go.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support on bare metal and virtualized platforms.
  • Design and productize AI cluster solutions (100k+ GPUs).
  • Develop automation for on-prem Kubernetes/AI clusters and OS.
  • Deploy core infra: databases, monitoring, and distributed storage.
  • Mentor junior engineers and lead the team to technical excellence.

Skills

Kubernetes
Linux
Bash/Python scripting
Go/Python/C++ development

Education

Bachelor’s degree in CS/IT/Engineering
7+ years experience without degree

Tools

Terraform
Ansible
OCI containers/Kubernetes

Job description

SpaceX is seeking a Sr. Site Reliability Engineer (Starshield) to design, operate, and scale GPU/CPU infrastructure for critical national security data centers in a Starshield environment.

You will automate on-prem Kubernetes/AI clusters, deploy databases and monitoring, and collaborate with AI engineers to deliver scalable, reliable software products. The role requires Top Secret/DOE-level clearance, 5+ years in Linux/Kubernetes, and strong scripting in Bash or Python.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI/GPU Infra & On-Prem Clusters
Senior SRE: AI/GPU Infra & On-Prem Clusters

SpaceX • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term incentives
Medical/Vision/Dental
+4
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Senior SRE for AI GPU Infra (Top Secret)
Senior SRE for AI GPU Infra (Top Secret)

SpaceX • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options and long-term incentives
Comprehensive medical/dental/vision
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
+4