Senior SRE for AI GPU Infra (Top Secret)

SpaceX

Redmond (WA)

On-site

USD 165,000 - 270,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options and long-term incentives
Comprehensive medical/dental/vision

Job summary

SpaceX in Redmond, WA is hiring a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale on-prem GPU/CPU infrastructure for a government-grade constellation.

You will automate deployments, develop scalable software, and mentor junior engineers, with Top Secret/SCI eligibility and extensive collaboration with AI teams. You will work on Kubernetes-based AI clusters, databases, and storage, driving reliability and performance in a highly secure environment.

Qualifications

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline.
  • 5+ years of professional experience with Linux operating systems or 7+ years in software/DevOps in lieu of a degree.
  • 5+ years of experience with Kubernetes.
  • Experience with Terraform, Ansible, or other infrastructure tools.
  • Experience with containerization technologies (OCI containers, Kubernetes).
  • Scripting in Bash, Python, or similar languages.
  • Development experience in Python, C++, or Go.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU as a service for external customers on bare metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions (100k+ GPU scale).
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters and OSes.
  • Deploy and manage core infrastructure such as databases, monitoring, and distributed storage.
  • Collaborate with AI engineers to create scalable, operable products.
  • Own lifecycle of services from inception to refinement.
  • Ensure high availability through monitoring and alerting.

Skills

Linux administration
Kubernetes
Python
C++
Go

Education

Bachelor's degree in computer science or related field

Tools

Terraform
Ansible
OCI containers/Kubernetes

Job description

SpaceX in Redmond, WA is hiring a Sr. Site Reliability Engineer (STARSHIELD) to design, operate, and scale on-prem GPU/CPU infrastructure for a government-grade constellation.

You will automate deployments, develop scalable software, and mentor junior engineers, with Top Secret/SCI eligibility and extensive collaboration with AI teams. You will work on Kubernetes-based AI clusters, databases, and storage, driving reliability and performance in a highly secure environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infrastructure & GPU Systems
Senior SRE: AI Infrastructure & GPU Systems

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
401(k) plan
Paid time off
+1
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI GPU Infra & On-Prem Systems
Senior SRE: AI GPU Infra & On-Prem Systems

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock options / long-term incentives
Health, vision & dental coverage
401(k) retirement plan
+3
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
AI Infrastructure SRE - GPU Clusters
AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

Hybrid
USD 125,000 - 200,000
Stock options
401(k) plan
Comprehensive medical benefits
+1
Site Reliability Engineer — AI GPU Infrastructure
Site Reliability Engineer — AI GPU Infrastructure

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision & dental benefits
401(k) retirement plan
+2