Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP

Hawthorne (CA)

On-site

USD 165,000 - 265,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
Long-term cash awards
Medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
Paid vacation
Paid holidays
Employee stock purchase plan

Job summary

SPACE EXPLORATION TECHNOLOGIES CORP in Hawthorne, CA is seeking a Sr. Site Reliability Engineer for Starshield. You will design, operate, and scale GPU and AI infrastructure, automation, and on‑premise Kubernetes clusters to support critical national security missions.

You will mentor juniors, lead technical excellence, collaborate with AI teams, and address reliability, performance, and security requirements. A Top Secret/DOE clearance is preferred, with travel and extended hours as needed.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go
  • 5+ years of experience with Python and Python-based development frameworks
  • Experience managing Kubernetes clusters, not just using them
  • Knowledge of Linux boot process and systems configuration
  • Deep understanding of testing, continuous integration, build, deployment & continuous monitoring
  • Understanding of relevant build technologies, such as Bazel and Makefiles
  • Focus on performance bottlenecks and performance improvement techniques
  • Understanding of distributed databases and data modeling
  • Experience with automatically managing thousands of servers (eg: Terraform or Ansible)
  • Strong networking knowledge of TCP/IP
  • Cloud virtualization experience

Responsibilities

  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions that enable high system availability
  • Mentor and train junior engineers
  • As a senior engineer you must lead the team to technical excellence – your decisions guide the team

Skills

Kubernetes
Linux
Terraform
Ansible
Python
Go
C++
Bash
GTX/GPUs
Networking TCP/IP

Education

Bachelor’s degree in CS/IT/Engineering

Tools

OCI containers
Kubernetes clusters management
CI/CD tooling

Job description

SPACE EXPLORATION TECHNOLOGIES CORP in Hawthorne, CA is seeking a Sr. Site Reliability Engineer for Starshield. You will design, operate, and scale GPU and AI infrastructure, automation, and on‑premise Kubernetes clusters to support critical national security missions.

You will mentor juniors, lead technical excellence, collaborate with AI teams, and address reliability, performance, and security requirements. A Top Secret/DOE clearance is preferred, with travel and extended hours as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI GPU Infra for Starshield & Space Missions
Senior SRE: AI GPU Infra for Starshield & Space Missions

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock incentives
Medical, vision and dental coverage
401(k) retirement plan
Senior SRE - AI GPU Infra, Kubernetes
Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen • Washington, Northern (KY)

Hybrid
USD 165,000 - 265,000
Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays
SRE: AI Infrastructure & GPU Platform Engineer
SRE: AI Infrastructure & GPU Platform Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE for AI Infrastructure & GPU Clusters
SRE for AI Infrastructure & GPU Clusters

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision, and dental coverage
401(k) plan
+4
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
Senior SRE – Starshield Infra, TS/SCI Clearance
Senior SRE – Starshield Infra, TS/SCI Clearance

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+2