SRE: AI GPU Infrastructure for Starshield

SpaceX

Redmond (WA)

On-site

USD 125,000 - 200,000

Full time

12 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical insurance
Vision insurance
Dental coverage
401(k) plan
Disability insurance
Life insurance
Parental leave
Employee stock purchase plan
Employee discounts

Job summary

SpaceX is seeking an experienced Site Reliability Engineer to design, operate, and scale on-premise GPU/CPU infrastructure for the Starshield program. You will deploy AI clusters at Very large scale and build automation to manage Kubernetes-based systems and OS environments.

You will collaborate with AI teams to deliver highly scalable and maintainable software products while ensuring high availability and secure access to classified environments.

Qualifications

  • Bachelor's degree or 3+ years of professional SRE/DevOps experience
  • 1+ years of professional Linux experience
  • Experience with infrastructure tools (Terraform, Ansible)
  • Experience with containerization (Kubernetes) and OCI containers
  • Scripting experience in Bash and Python; development in Python/C++/Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Provide GPU-as-a-service for external customers on bare metal and virtualized platforms
  • Design, validate and productize AI cluster solutions (100k+ GPUs)
  • Develop automation for on-premise Kubernetes/AI clusters and OS management
  • Deploy and manage core infrastructure (databases, monitoring, storage)
  • Collaborate with AI engineers on scalable, operable products
  • Oversee full lifecycle of services from design to deployment and refinement
  • Implement monitoring/alerting for high availability
  • Identify improvements and create innovative solutions for reliability

Skills

Linux
Terraform
Ansible
Kubernetes
Python
Go
C++
Bash
Top Secret clearance
TCP/IP networking

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Ansible
Kubernetes

Job description

SpaceX is seeking an experienced Site Reliability Engineer to design, operate, and scale on-premise GPU/CPU infrastructure for the Starshield program. You will deploy AI clusters at Very large scale and build automation to manage Kubernetes-based systems and OS environments.

You will collaborate with AI teams to deliver highly scalable and maintainable software products while ensuring high availability and secure access to classified environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
SRE for AI Infrastructure & GPU Clusters
SRE for AI Infrastructure & GPU Clusters

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision, and dental coverage
401(k) plan
+4
Senior SRE - AI GPU Infra, Kubernetes
Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen • Washington, Northern (KY)

Hybrid
USD 165,000 - 265,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
Senior SRE: AI GPU Infra for Starshield & Space Missions
Senior SRE: AI GPU Infra for Starshield & Space Missions

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock incentives
Medical, vision and dental coverage
401(k) retirement plan
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Senior SRE — Starshield Cloud & Reliability
Senior SRE — Starshield Cloud & Reliability

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+1
Senior SRE: Secure, Scalable Infra & Automation
Senior SRE: Secure, Scalable Infra & Automation

SpaceX • Washington

On-site
USD 165,000 - 230,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+1