AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP

Redmond, Northern (WA, KY)

Hybrid

USD 125,000 - 200,000

Full time

20 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) plan
Comprehensive medical benefits
Paid vacation and holidays

Job summary

SpaceX is seeking a Site Reliability Engineer for AI Infrastructure (Starshield) in Redmond, WA. The role focuses on designing, operating, and scaling GPU-backed AI infrastructure and on-premise Kubernetes clusters to support national security missions.

You will collaborate with AI engineers, develop automation, and ensure high availability across systems while navigating sensitive security requirements.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 1+ years of professional SRE/DevOps experience or 3+ years without a degree.
  • 1+ years of professional Linux experience.
  • Experience with infrastructure tools such as Terraform, Ansible, or similar.
  • Experience with containerization and Kubernetes.
  • Scripting experience with Bash, Python, or similar.
  • Development experience in Python, C++, or Go.

Responsibilities

  • Manage GPU service for external customers on bare metal and virtualized platforms.
  • Design and productize AI cluster solutions at large scale (100k+ GPUs).
  • Develop automation to deploy and manage on-premise KubernetesAI clusters and OSes.
  • Deploy and manage core infrastructure including databases, monitoring, and storage.
  • Collaborate with AI engineers to build scalable, operable products.
  • Oversee full lifecycle of services from design to operation and refinement.
  • Support monitoring and alerting to ensure high availability.
  • Identify improvement areas and implement innovative solutions for reliability.

Skills

Linux
Bash
Python
C++
Go
DevOps
Kubernetes
Networking
Cloud computing

Education

Bachelor's degree in CS/IT/Engineering

Tools

Terraform
Ansible
Docker
Kubernetes

Job description

SpaceX is seeking a Site Reliability Engineer for AI Infrastructure (Starshield) in Redmond, WA. The role focuses on designing, operating, and scaling GPU-backed AI infrastructure and on-premise Kubernetes clusters to support national security missions.

You will collaborate with AI engineers, develop automation, and ensure high availability across systems while navigating sensitive security requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure SRE — GPU & On-Prem Kubernetes
AI Infrastructure SRE — GPU & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Washington

On-site
USD 125,000 - 195,000
Stock options/long-term incentives
Medical, Vision, Dental
401(k) plan
+2
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
Site Reliability Engineer — AI Infra & GPU Clusters
Site Reliability Engineer — AI Infra & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k)
Health coverage
+1
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE - AI GPU Infra, Kubernetes
Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen • Washington, Northern (KY)

Hybrid
USD 165,000 - 265,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1