SRE, AI GPU Infra for Starshield

SpaceX

Palo Alto (CA)

On-site

USD 125,000 - 195,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options / equity
Medical, vision, dental
401(k)
Paid parental leave
3 weeks paid vacation
10+ paid holidays
Sick leave
Potential signing bonus

Job summary

SpaceX is hiring a Site Reliability Engineer (STARSHIELD) to design, operate, and scale on-premise GPU/CPU infrastructure for critical national security missions. You will automate deployments, build scalable software, and collaborate with AI teams across the company.

The role requires Linux expertise, Terraform/Ansible, and container tech like Kubernetes, with a TS/SCI or DOE-level clearance considered. SpaceX offers strong compensation and benefits.

Qualifications

  • Bachelor’s degree in CS/IT or engineering with 1+ year in SRE/DevOps; or 3+ years in SRE/DevOps without a degree.
  • 1+ years of Linux experience and scripting in Bash, Python or similar.
  • Experience with Terraform, Ansible or other infra tools.
  • Experience with containerization (Kubernetes, OCI).
  • Proficient in Python, C++, or Go development.

Responsibilities

  • Manage GPU/CPU infrastructure deployments in Top Secret data centers.
  • Provide support for GPU as a service on bare metal and virtualized platforms.
  • Design and productize AI cluster solutions at 100k+ GPU scale.
  • Automate deployment and management of on-prem Kubernetes/AI clusters and OSes.
  • Deploy core infra: databases, monitoring, and distributed storage.
  • Collaborate with AI engineers for scalable, operable products.
  • Improve the full lifecycle of services from design to operation.
  • Implement monitoring and alerting for high availability.
  • Identify improvements and develop innovative solutions for reliability.

Skills

Linux
Python
Go
C++
Bash
DevOps
SRE
Cloud
Networking

Education

Bachelor’s degree in CS/IT/Engineering

Tools

Terraform
Ansible
Kubernetes
OCI containers
Docker
Python tooling

Job description

SpaceX is hiring a Site Reliability Engineer (STARSHIELD) to design, operate, and scale on-premise GPU/CPU infrastructure for critical national security missions. You will automate deployments, build scalable software, and collaborate with AI teams across the company.

The role requires Linux expertise, Terraform/Ansible, and container tech like Kubernetes, with a TS/SCI or DOE-level clearance considered. SpaceX offers strong compensation and benefits.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3
SRE: AI GPU Infrastructure & On-Prem Kubernetes
SRE: AI GPU Infrastructure & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Performance bonuses
Comprehensive health coverage (medical
+2
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
Senior SRE – AI GPU Infra for Global Missions
Senior SRE – AI GPU Infra for Global Missions

InvestedintheMission • Washington

On-site
USD 165,000 - 265,000
Stock options
Employee Stock Purchase Plan
Medical, vision, and dental coverage
+3
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
+4
Senior SRE: AI Infrastructure & GPU Systems
Senior SRE: AI Infrastructure & GPU Systems

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
401(k) plan
Paid time off
+1
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
401(k) plan
Medical, vision, dental coverage
+2
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6