Senior SRE: AI Infra & GPU Clusters

SpaceX

Washington (District of Columbia)

On-site

USD 165,000 - 265,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

SpaceX is seeking a Sr. Site Reliability Engineer for the Starshield program to design, operate, and scale on-premise GPU/CPU infrastructure and AI clusters.

You will develop automation to deploy Kubernetes and related OSes, deploy databases and monitoring, and collaborate with AI engineers to deliver reliable services. The role requires deep Linux and Kubernetes expertise, Terraform/Ansible experience, and the ability to lead a team toward technical excellence while enabling high availability

Qualifications

  • Bachelor’s degree and 5+ years Linux or 7+ years software/DevOps/SRE in lieu of degree
  • 5+ years Kubernetes experience
  • 5+ years managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (OCI containers, Kubernetes)
  • Scripting in Bash, Python, or similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Collaborate with AI engineers to create scalable, operable products
  • Lead the lifecycle of services from inception to refinement
  • Mentor and train junior engineers

Skills

Linux
Kubernetes
Terraform
Ansible
Bash
Python
C++
Go
Networking
Cloud

Education

Bachelor's degree in CS/IT or engineering

Tools

Docker
OCI containers
NVIDIA GPU deployment stacks

Job description

SpaceX is seeking a Sr. Site Reliability Engineer for the Starshield program to design, operate, and scale on-premise GPU/CPU infrastructure and AI clusters.

You will develop automation to deploy Kubernetes and related OSes, deploy databases and monitoring, and collaborate with AI engineers to deliver reliable services. The role requires deep Linux and Kubernetes expertise, Terraform/Ansible experience, and the ability to lead a team toward technical excellence while enabling high availability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3
SRE: AI GPU Infra & On-Prem Kubernetes
SRE: AI GPU Infra & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k) retirement plan
+2
AI Infrastructure SRE - GPU Clusters
AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

Hybrid
USD 125,000 - 200,000
Stock options
401(k) plan
Comprehensive medical benefits
+1
Senior SRE: AI GPU Infra & On-Prem Systems
Senior SRE: AI GPU Infra & On-Prem Systems

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock options / long-term incentives
Health, vision & dental coverage
401(k) retirement plan
+3
Senior SRE: AI Infrastructure & GPU Systems
Senior SRE: AI Infrastructure & GPU Systems

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
401(k) plan
Paid time off
+1