Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission

Palo Alto (CA)

On-site

USD 165,000 - 265,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
Paid vacation (3 weeks)
Paid holidays (10+)
Sick leave

Job summary

SpaceX is seeking a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale on‑premise GPU and AI infrastructure for a national security‑focused constellation.

The role emphasizes automation, Kubernetes, and cross‑team collaboration with AI engineers to deliver highly available services. You will mentor junior engineers, lead critical infra initiatives, and work across data centers with stringent security requirements, including Top Secret clearance considerations.

Qualifications

  • Bachelors or higher in CS/IT/Engineering or equivalent with 5+ years of Linux/SRE/DevOps experience.
  • 5+ years of Kubernetes administration and 5+ years of Linux OS management.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide support for GPU as a service on bare metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions (100k+ GPU scale).
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters and OSes.
  • Deploy core infrastructure including databases, monitoring, and distributed storage.
  • Collaborate with AI engineers to build scalable, operable products.
  • Oversee the full lifecycle of services from design to operation.
  • Maintain high availability with robust monitoring and alerting.
  • Identify improvement opportunities and implement innovative solutions.
  • Mentor and train junior engineers; lead the team to technical excellence.

Skills

Kubernetes
Linux
Terraform
Ansible
Python
C++
Go
Bash

Education

Bachelor's degree in Computer Science, Information Systems, or Engineering

Tools

Git
CI/CD tools

Job description

SpaceX is seeking a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale on‑premise GPU and AI infrastructure for a national security‑focused constellation.

The role emphasizes automation, Kubernetes, and cross‑team collaboration with AI engineers to deliver highly available services. You will mentor junior engineers, lead critical infra initiatives, and work across data centers with stringent security requirements, including Top Secret clearance considerations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
SRE: AI GPU Infrastructure & On-Prem Kubernetes
SRE: AI GPU Infrastructure & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Performance bonuses
Comprehensive health coverage (medical
+2
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE – AI GPU Infra for Global Missions
Senior SRE – AI GPU Infra for Global Missions

InvestedintheMission • Washington

On-site
USD 165,000 - 265,000
Stock options
Employee Stock Purchase Plan
Medical, vision, and dental coverage
+3
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
401(k) plan
Medical, vision, dental coverage
+2
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
SRE: AI GPU Infra & On-Prem Kubernetes
SRE: AI GPU Infra & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k) retirement plan
+2
Senior SRE: AI/GPU Infra & On-Prem Clusters
Senior SRE: AI/GPU Infra & On-Prem Clusters

SpaceX • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term incentives
Medical/Vision/Dental
+4
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000