Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX

Union Hill-Novelty Hill (WA)

On-site

USD 165,000 - 270,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Company stock
Long-term incentives
Employee Stock Purchase Plan
Medical, vision, dental
401(k) plan
Paid parental leave
Paid vacation
10+ holidays per year
Shuttle service

Job summary

SpaceX is seeking a Sr. Site Reliability Engineer (Starshield) to design, operate and scale GPU/CPU infrastructure for the Starshield constellation.

You will automate on-prem resources, deploy Kubernetes/AI clusters, and collaborate with AI engineers to deliver scalable, reliable systems. Requirements include 5+ years in Linux environments and Kubernetes, Terraform/Ansible tools, and programming in Python/Go/C++.

Qualifications

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions that enable high system availability
  • Mentor and train junior engineers
  • As a senior engineer you must lead the team to technical excellence - your decisions guide the team

Skills

Kubernetes experience
Linux administration
Scripting (Bash/Python)
Programming (Python/Go/C++)
Communication skills

Education

Bachelor's degree in computer science / information systems / engineering

Tools

Terraform
Ansible
OCI containers
Kubernetes (management)

Job description

SpaceX is seeking a Sr. Site Reliability Engineer (Starshield) to design, operate and scale GPU/CPU infrastructure for the Starshield constellation.

You will automate on-prem resources, deploy Kubernetes/AI clusters, and collaborate with AI engineers to deliver scalable, reliable systems. Requirements include 5+ years in Linux environments and Kubernetes, Terraform/Ansible tools, and programming in Python/Go/C++.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive benefits package
Paid time off and holidays
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
401(k) plan
Medical, vision, dental coverage
+2
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI/GPU Infra & On-Prem Clusters
Senior SRE: AI/GPU Infra & On-Prem Clusters

SpaceX • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term incentives
Medical/Vision/Dental
+4
SRE: AI GPU Infra & On-Prem Kubernetes
SRE: AI GPU Infra & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k) retirement plan
+2
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes
Senior AI Infra SRE — GPU Clusters & On-Prem Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Stock options (long-term incentives)
Medical, vision, dental coverage
401(k) retirement plan
+4
SRE: AI GPU Infrastructure & On-Prem Kubernetes
SRE: AI GPU Infrastructure & On-Prem Kubernetes

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
Performance bonuses
Comprehensive health coverage (medical
+2