SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX

Union Hill-Novelty Hill (WA)

On-site

USD 125,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
Paid time off
Paid holidays
Vacation and parental leave
Employee shuttle program

Job summary

SpaceX seeks a Site Reliability Engineer (Starshield) to design, operate, and scale on-prem GPU/CPU infrastructure supporting a national security mission. You will automate deployments, build scalable software, and work with AI teams to deliver reliable systems.

You will manage Kubernetes clusters, cloud-like on‑prem environments, and performance monitoring while coordinating with engineers across teams. This role requires willingness to extended hours and travel as needed.

Qualifications

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline with 1+ years of SRE/DevOps; OR 3+ years in lieu of a degree.
  • 1+ years of experience with Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Manage and support GPU as a service for external customers on bare metal and virtualized platforms
  • Design, validate, and productize AI cluster solutions (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters and OSes
  • Deploy and manage databases, monitoring, and distributed storage
  • Collaborate with AI engineers to create scalable, operable products
  • Improve services lifecycle from design to operation and refinement
  • Maintain high availability through monitoring and alerting
  • Identify and implement performance improvements

Skills

Linux
Terraform
Ansible
Kubernetes
Bash
Python
C++
Go

Education

Bachelor's degree in CS/IT/Engineering
3+ years SRE/DevOps experience (in lieu of degree)

Tools

OCI containers
Git
CI/CD tooling

Job description

SpaceX seeks a Site Reliability Engineer (Starshield) to design, operate, and scale on-prem GPU/CPU infrastructure supporting a national security mission. You will automate deployments, build scalable software, and work with AI teams to deliver reliable systems.

You will manage Kubernetes clusters, cloud-like on‑prem environments, and performance monitoring while coordinating with engineers across teams. This role requires willingness to extended hours and travel as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE: AI Infrastructure & GPU Platform Engineer
SRE: AI Infrastructure & GPU Platform Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Senior SRE - AI GPU Infra, Kubernetes
Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen • Washington, Northern (KY)

Hybrid
USD 165,000 - 265,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE for AI Infrastructure & GPU Clusters
SRE for AI Infrastructure & GPU Clusters

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision, and dental coverage
401(k) plan
+4
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
Senior SRE: AI GPU Infra for Starshield & Space Missions
Senior SRE: AI GPU Infra for Starshield & Space Missions

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock incentives
Medical, vision and dental coverage
401(k) retirement plan
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Senior SRE — Starshield Cloud & Reliability
Senior SRE — Starshield Cloud & Reliability

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+1