Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP

Palo Alto (CA)

On-site

USD 165,000 - 265,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock options
Comprehensive medical, vision & dental
Paid vacation & holidays

Job summary

SpaceX is hiring a Sr. Site Reliability Engineer focused on AI infrastructure for Starshield. You will design, operate, and scale GPU-enabled on-premise and Kubernetes-based AI clusters, and automate deployment and management of core services, databases, and storage.

You will mentor junior engineers, drive reliability across the lifecycle, and collaborate with AI teams to deliver scalable, operable products while supporting 100k+ GPU scale and demanding security clearances.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years in lieu of a degree.
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions that enable high system availability
  • Mentor and train junior engineers
  • As a senior engineer you must lead the team to technical excellence – your decisions guide the team

Skills

Kubernetes
Linux
Terraform
Ansible
Python
Clearance

Education

Bachelor’s degree in computer science, information systems/IT, or engineering

Tools

OCI containers
Terraform
Ansible

Job description

SpaceX is hiring a Sr. Site Reliability Engineer focused on AI infrastructure for Starshield. You will design, operate, and scale GPU-enabled on-premise and Kubernetes-based AI clusters, and automate deployment and management of core services, databases, and storage.

You will mentor junior engineers, drive reliability across the lifecycle, and collaborate with AI teams to deliver scalable, operable products while supporting 100k+ GPU scale and demanding security clearances.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI Infra & GPU Clusters
Senior SRE: AI Infra & GPU Clusters

SpaceX • Washington

On-site
USD 165,000 - 265,000
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
Senior SRE: AI GPU Infra & On-Prem Kubernetes
Senior SRE: AI GPU Infra & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

On-site
USD 165,000 - 270,000
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
AI Infrastructure SRE - GPU Clusters
AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

Hybrid
USD 125,000 - 200,000
Stock options
401(k) plan
Comprehensive medical benefits
+1
SRE for AI GPU Infrastructure
SRE for AI GPU Infrastructure

SPACE EXPLORATION TECHNOLOGIES CORP • Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
401(k)
Medical, vision, dental
+3
Senior SRE: AI Infrastructure & GPU Systems
Senior SRE: AI Infrastructure & GPU Systems

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
401(k) plan
Paid time off
+1
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3