Site Reliability Engineer - AI GPU Infrastructure

Spacex

Washington (District of Columbia)

On-site

USD 125,000 - 195,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Long-term incentives
Health, vision and dental coverage
401(k) retirement plan
Paid parental leave
3 weeks paid vacation and 10+ holidays

Job summary

SpaceX is seeking a Site Reliability Engineer to design, operate, and scale Starshield software and GPU infrastructure supporting critical national security missions. You will develop automation to deploy on-premise resources, build scalable AI clusters, and collaborate with AI engineers across teams.

The role covers lifecycle management, high-availability systems, and operations across Top Secret datacenters, with a focus on performance, reliability, and security in a government-facing program.

Qualifications

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline with 1+ years in SRE/DevOps; or 3+ years in lieu of a degree.
  • 1+ years of professional experience with Linux operating systems.
  • Experience with Terraform, Ansible, or other infrastructure tools.
  • Experience with containerization technologies (Kubernetes).
  • Experience scripting in Bash, Python, or similar languages.
  • Development experience in Python, C++, or Go.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters.
  • Provide support for GPU as a service on bare metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions (100k+ GPU scale).
  • Develop automation to deploy and manage on-premise Kubernetes AI clusters and OSes.
  • Deploy and manage core infrastructure such as databases, monitoring, and distributed storage.
  • Collaborate with AI engineers to create scalable, operable products.
  • Improve lifecycle of services from inception to refinement.
  • Ensure high availability through monitoring and alerting.
  • Identify areas for improvement and create innovative solutions.

Skills

Linux
Terraform
Ansible
Kubernetes
Bash
Python
Go
C++

Education

Bachelor's degree in computer science or engineering

Tools

Kubernetes
Terraform
Ansible

Job description

SpaceX is seeking a Site Reliability Engineer to design, operate, and scale Starshield software and GPU infrastructure supporting critical national security missions. You will develop automation to deploy on-premise resources, build scalable AI clusters, and collaborate with AI engineers across teams.

The role covers lifecycle management, high-availability systems, and operations across Top Secret datacenters, with a focus on performance, reliability, and security in a government-facing program.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: AI Infrastructure & GPU Platform Engineer
SRE: AI Infrastructure & GPU Platform Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE for AI Infrastructure & GPU Clusters
SRE for AI Infrastructure & GPU Clusters

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision, and dental coverage
401(k) plan
+4
Senior SRE - AI GPU Infra, Kubernetes
Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen • Washington, Northern (KY)

Hybrid
USD 165,000 - 265,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI GPU Infra for Starshield & Space Missions
Senior SRE: AI GPU Infra for Starshield & Space Missions

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock incentives
Medical, vision and dental coverage
401(k) retirement plan
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
Site Reliability Engineer, AI Infrastructure (Starshield)
Site Reliability Engineer, AI Infrastructure (Starshield)

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3