Senior SRE - AI GPU Infra, Kubernetes

Artha Nexgen

Washington, Northern (District of Columbia, KY)

Hybrid

USD 165,000 - 265,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, vision, and dental coverage
401(k) retirement plan
Paid vacation and holidays

Job summary

SpaceX is seeking a Sr. Site Reliability Engineer to design, operate, and scale the Starshield AI infrastructure. You will build automation for on‑prem Kubernetes, deploy GPU clusters, and collaborate with AI teams to deliver highly scalable and reliable services.

The role requires senior technical leadership, experience with Linux, Kubernetes, Terraform/Ansible, and strong Python/C++/Go development. International travel and security clearances may apply.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline with 5+ years on Linux; or 7+ years in software/DevOps/SRE in lieu of a degree.
  • 5+ years with Kubernetes and Linux.
  • Experience with Terraform, Ansible or other infra tools; containerization with OCI/Kubernetes.
  • Scripting in Bash, Python, or similar; development in Python, C++, or Go.

Responsibilities

  • Manage GPU as a service for external customers on bare metal and virtualized platforms.
  • Design, validate and productize AI cluster solutions at scale.
  • Develop automation for on‑prem Kubernetes/AI clusters and OS deployment.
  • Deploy core infra: databases, monitoring, and distributed storage.
  • Collaborate with AI engineers to build scalable, operable products.

Skills

Linux
Kubernetes
Python
C++
Go
Bash scripting
Networking TCP/IP

Education

Bachelor’s degree in CS/Engineering

Tools

Terraform
Ansible
Kubernetes (cluster management)

Job description

SpaceX is seeking a Sr. Site Reliability Engineer to design, operate, and scale the Starshield AI infrastructure. You will build automation for on‑prem Kubernetes, deploy GPU clusters, and collaborate with AI teams to deliver highly scalable and reliable services.

The role requires senior technical leadership, experience with Linux, Kubernetes, Terraform/Ansible, and strong Python/C++/Go development. International travel and security clearances may apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI-GPU Infra & On-Prem Kubernetes
Senior SRE: AI-GPU Infra & On-Prem Kubernetes

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 165,000 - 270,000
Company stock
Long-term incentives
Employee Stock Purchase Plan
+6
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4
Senior SRE: AI GPU Infra for Starshield & Space Missions
Senior SRE: AI GPU Infra for Starshield & Space Missions

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock incentives
Medical, vision and dental coverage
401(k) retirement plan
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
SRE: AI Infrastructure & GPU Platform Engineer
SRE: AI Infrastructure & GPU Platform Engineer

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Senior SRE: AI Infrastructure & GPU Clusters
Senior SRE: AI Infrastructure & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term cash awards
Medical, vision, dental coverage
+5
SRE for AI Infrastructure & GPU Clusters
SRE for AI Infrastructure & GPU Clusters

United States Digital Space LLC • United States

Remote
USD 125,000 - 195,000
Stock options
Medical, vision, and dental coverage
401(k) plan
+4
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
Senior SRE — Starshield Cloud & Reliability
Senior SRE — Starshield Cloud & Reliability

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 230,000
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
Paid parental leave
+1