SRE: AI GPU Infrastructure & On-Prem Kubernetes

InvestedintheMission

Hawthorne (CA)

On-site

USD 125,000 - 195,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Performance bonuses
Comprehensive health coverage (medical
401(k) retirement plan
Paid time off

Job summary

SpaceX is recruiting a Site Reliability Engineer (STARSHIELD) to design, operate, and scale infrastructure supporting critical national security missions. You will build automation for on‑premise compute resources, deploy AI clusters at scale, and collaborate across engineering teams to ensure highly available services.

You will work with Linux, Terraform, Ansible, and Kubernetes, with responsibilities spanning from deployment to monitoring and lifecycle refinement.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline; 1+ years in SRE/DevOps or 3+ years in lieu of a degree
  • 1+ years of professional experience with Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (OCI containers, Kubernetes)
  • Scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Collaborate with AI engineers to create scalable, operable products
  • Engage in and improve the lifecycle of services from inception to refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions for high availability

Skills

Linux
Terraform
Ansible
Kubernetes
Bash
Python
Go
C++

Education

Bachelor’s degree in CS/IS/Engineering

Tools

OCI containers
Kubernetes
Terraform
Ansible

Job description

SpaceX is recruiting a Site Reliability Engineer (STARSHIELD) to design, operate, and scale infrastructure supporting critical national security missions. You will build automation for on‑premise compute resources, deploy AI clusters at scale, and collaborate across engineering teams to ensure highly available services.

You will work with Linux, Terraform, Ansible, and Kubernetes, with responsibilities spanning from deployment to monitoring and lifecycle refinement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
+4
SRE: AI GPU Infra & On-Prem Kubernetes
SRE: AI GPU Infra & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k) retirement plan
+2
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE, AI GPU Infra for Starshield
SRE, AI GPU Infra for Starshield

SpaceX • Palo Alto (CA)

On-site
USD 125,000 - 195,000
Stock options / equity
Medical, vision, dental
401(k)
+5
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000