Senior SRE: AI Infra & GPU Clusters

SpaceX

Washington

On-site

USD 165,000 - 265,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

SpaceX is hiring a Senior Site Reliability Engineer for Starshield to design, operate, and scale GPU/CPU infrastructure in secure data centers. You will develop automation for on-prem Kubernetes/AI clusters and collaborate with AI teams to deliver scalable, maintainable software products.

You will mentor engineers, ensure high availability, and guide system design with an emphasis on reliability and performance in a defense-related program.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or engineering with 5+ years on Linux OR 7+ years in software/DevOps/SRE without a degree.
  • 5+ years of Kubernetes experience.
  • 5+ years of Linux OS experience.
  • Experience with Terraform, Ansible, or other infrastructure tools.
  • Experience with containerization (OCI, Kubernetes).
  • Scripting in Bash, Python, or similar languages.
  • Development experience in Python, C++, or Go.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support on bare metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions (100k+ GPU).
  • Develop automation to deploy and manage on-prem Kubernetes/AI clusters and OSes.
  • Deploy and manage core infra: databases, monitoring, storage.
  • Collaborate with AI engineers to build scalable, operable products.
  • Lead lifecycle from design to deployment and refinement.
  • Ensure high availability through monitoring and alerting.
  • Identify improvements and implement innovative solutions.
  • Mentor and train junior engineers.

Skills

Kubernetes
Linux
Terraform
Ansible
Containerization
Python
Go/C++
Bash scripting
DevOps

Education

Bachelor's degree in Computer Science or related field
7+ years experience in lieu of degree

Tools

Kubernetes
Terraform
Ansible

Job description

SpaceX is hiring a Senior Site Reliability Engineer for Starshield to design, operate, and scale GPU/CPU infrastructure in secure data centers. You will develop automation for on-prem Kubernetes/AI clusters and collaborate with AI teams to deliver scalable, maintainable software products.

You will mentor engineers, ensure high availability, and guide system design with an emphasis on reliability and performance in a defense-related program.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
AI Infrastructure SRE - GPU Clusters
AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

Hybrid
USD 125,000 - 200,000
Stock options
401(k) plan
Comprehensive medical benefits
+1
Senior SRE: AI Infrastructure & GPU Systems
Senior SRE: AI Infrastructure & GPU Systems

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
401(k) plan
Paid time off
+1
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3
SRE: AI GPU Infra & On-Prem Kubernetes
SRE: AI GPU Infra & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k) retirement plan
+2