Senior SRE – AI GPU Infra for Global Missions

InvestedintheMission

Washington (District of Columbia)

On-site

USD 165,000 - 265,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options
Employee Stock Purchase Plan
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
Sick leave

Job summary

SpaceX seeks a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale the infrastructure behind Starshield's AI and GPU platforms.

You will automate on‑premise compute resources, build scalable software, and collaborate with broader engineering teams to ensure reliability of critical national security missions. The role covers SRE/DevOps for GPU clusters, Kubernetes, databases, and storage, with opportunities to mentor engineers and lead technical excellence.

Qualifications

  • Bachelors in CS/IT/engineering or equivalent with 5+ years Linux experience.
  • 5+ years Kubernetes experience.
  • 5+ years Linux OS management.
  • Terraform, Ansible or similar tools.
  • Containerization tech knowledge (OCI, Kubernetes).
  • Scripting in Bash/Python or similar.
  • Development in Python/C++/Go.

Responsibilities

  • Deploy and manage GPU/CPU infrastructure for Top Secret data centers.
  • Provide GPU-as-a-service on bare metal and virtualized platforms.
  • Design and productize AI cluster solutions (100k+ GPUs).
  • Develop automation for on-prem Kubernetes/AI clusters and OS management.
  • Deploy core infra: databases, monitoring, and distributed storage.
  • Collaborate with AI engineers on scalable, operable products.
  • Own full lifecycle of services from inception to refinement.
  • Implement monitoring and alerting for high availability.
  • Identify improvements and implement high-availability solutions.
  • Mentor and train junior engineers.
  • Lead the team to technical excellence as a senior engineer.

Skills

Kubernetes
Linux administration
Terraform
Ansible
Containerization
Python
C++
Go
Shell scripting
Security clearance

Education

Bachelor's degree in CS/IT/engineering

Tools

OCI containers
Kubernetes (AI clusters)

Job description

SpaceX seeks a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale the infrastructure behind Starshield's AI and GPU platforms.

You will automate on‑premise compute resources, build scalable software, and collaborate with broader engineering teams to ensure reliability of critical national security missions. The role covers SRE/DevOps for GPU clusters, Kubernetes, databases, and storage, with opportunities to mentor engineers and lead technical excellence.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infrastructure & GPU Systems
Senior SRE: AI Infrastructure & GPU Systems

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
401(k) plan
Paid time off
+1
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
+4
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
401(k) plan
Medical, vision, dental coverage
+2
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE, AI GPU Infra for Starshield
SRE, AI GPU Infra for Starshield

SpaceX • Palo Alto (CA)

On-site
USD 125,000 - 195,000
Stock options / equity
Medical, vision, dental
401(k)
+5
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4