SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP

Northern (KY)

Hybrid

USD 125,000 - 195,000

Full time

27 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
401(k)
Medical, vision, dental
Paid parental leave
Vacation and holidays
Sick leave

Job summary

SpaceX is seeking an experienced Site Reliability Engineer to design, operate, and scale AI infrastructure for Starshield. You will manage GPU services on bare metal and virtual platforms, build automation for on-prem Kubernetes/AI clusters, and collaborate across teams to ensure highly available systems.

The role demands strong Linux, Python/C++, and Terraform/Ansible experience, with security clearances and travel as needed. SpaceX offers a comprehensive benefits package.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline with 1+ years of professional SRE/DevOps experience; or 3+ years of experience in lieu of a degree.
  • 1+ years of professional experience with Linux operating systems.
  • Experience with Terraform, Ansible, or other infrastructure tools.
  • Experience with containerization technologies (e.g. OCI containers, Kubernetes).
  • Experience scripting in Bash, Python, or other similar languages.
  • Development experience in Python, C++, or Go.

Responsibilities

  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms.
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale).
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems.
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage.
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products.
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement.
  • Monitoring and alerting supporting systems to have high availability.
  • Identify areas for improvement and create innovative solutions that enable high system availability.

Skills

GPU
Terraform
Ansible
Kubernetes
Linux
Bash
Python
C++
Go
Networking

Education

Bachelor’s degree in computer science/engineering

Tools

Terraform
Ansible
Kubernetes

Job description

SpaceX is seeking an experienced Site Reliability Engineer to design, operate, and scale AI infrastructure for Starshield. You will manage GPU services on bare metal and virtual platforms, build automation for on-prem Kubernetes/AI clusters, and collaborate across teams to ensure highly available systems.

The role demands strong Linux, Python/C++, and Terraform/Ansible experience, with security clearances and travel as needed. SpaceX offers a comprehensive benefits package.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
Senior SRE - AI GPU Infra, Kubernetes
Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen • Washington, Northern (KY)

Hybrid
USD 165,000 - 265,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
AI Infrastructure SRE - GPU Clusters
AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

Hybrid
USD 125,000 - 200,000
Stock options
401(k) plan
Comprehensive medical benefits
+1
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
AI Infrastructure SRE — GPU & On-Prem Kubernetes
AI Infrastructure SRE — GPU & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Washington

On-site
USD 125,000 - 195,000
Stock options/long-term incentives
Medical, Vision, Dental
401(k) plan
+2
SRE for AI Infrastructure & GPU Clusters
SRE for AI Infrastructure & GPU Clusters

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision, and dental coverage
401(k) plan
+4
Site Reliability Engineer — AI Infra & GPU Clusters
Site Reliability Engineer — AI Infra & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k)
Health coverage
+1