AI Infra Engineer: Scale GPU Clusters & On-Prem AI

Xapply

Palo Alto (CA)

On-site

USD 125,000 - 195,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Long-term incentives
Medical, vision & dental
401(k) retirement plan
Paid parental leave
3 weeks vacation
10+ paid holidays
Sick leave

Job summary

SpaceX is seeking a Software Engineer for AI Infra (Starshield) to design, deploy, and scale GPU/CPU infrastructure across Top Secret data centers. You will automate on-premise Kubernetes AI clusters and supporting OS environments, while collaborating closely with AI engineers to deliver highly scalable products.

The role covers Site Reliability Engineering, DevOps, and GPU platforms with extensive automation, monitoring, and performance optimization.

Qualifications

  • Bachelor’s degree in a computing/engineering field or equivalent experience.
  • 1+ years of professional SRE/DevOps; or 3+ years without a degree.
  • Experience with Linux systems and scripting (Bash, Python, etc.).
  • Hands-on with Terraform, Ansible, Kubernetes; containerized apps.

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support on bare metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions at scale.
  • Develop automation for on-premise Kubernetes AI clusters and OSs.
  • Deploy and manage databases, monitoring, and distributed storage.
  • Collaborate with AI engineers to build scalable, maintainable products.
  • Oversee full service lifecycle from design to operation and refinement.
  • Maintain monitoring and alerting for high availability.
  • Identify improvements and implement innovative high-availability solutions.

Skills

SRE/DevOps
Linux
Python
Go
C++
Bash
Troubleshooting
Performance optimization

Education

Bachelor's degree in CS or related field

Tools

Terraform
Ansible
Kubernetes
OCI containers

Job description

SpaceX is seeking a Software Engineer for AI Infra (Starshield) to design, deploy, and scale GPU/CPU infrastructure across Top Secret data centers. You will automate on-premise Kubernetes AI clusters and supporting OS environments, while collaborating closely with AI engineers to deliver highly scalable products.

The role covers Site Reliability Engineering, DevOps, and GPU platforms with extensive automation, monitoring, and performance optimization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer: GPU Clusters & On‑Prem Systems
AI Infra Engineer: GPU Clusters & On‑Prem Systems

Spacex • Hawthorne (CA), Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k)
+1
AI Infrastructure Engineer: GPU, On‑Prem & Kubernetes
AI Infrastructure Engineer: GPU, On‑Prem & Kubernetes

Engg • Redmond (WA)

On-site
USD 125,000 - 200,000
Stock options
Medical, vision & dental coverage
401(k) plan
+2
AI Infrastructure Engineer — GPU & On-Prem Kubernetes
AI Infrastructure Engineer — GPU & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 190,000
Stock options and long‑term incentives
Discretionary bonuses
Comprehensive medical, vision, dental
Senior AI Infra Engineer: GPU-Driven On-Prem AI Clusters
Senior AI Infra Engineer: GPU-Driven On-Prem AI Clusters

InvestedintheMission • Washington

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical/vision/dental
Paid time off and holidays
AI Infrastructure Engineer - GPU & On-Prem
AI Infrastructure Engineer - GPU & On-Prem

InvestedintheMission • Washington

On-site
USD 125,000 - 195,000
Stock options
Bonuses
Medical, vision, dental coverage
+3
Senior AI Infrastructure Engineer — GPU & On-Prem Clusters
Senior AI Infrastructure Engineer — GPU & On-Prem Clusters

Xapply • Washington

On-site
USD 165,000 - 265,000
Stock options
Medical coverage
401(k)
+1
Senior AI Infra Engineer: GPU & Kubernetes Scale Leader
Senior AI Infra Engineer: GPU & Kubernetes Scale Leader

Engg • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Health insurance
Dental coverage
Vision coverage
+7
Senior AI Infrastructure Engineer - GPU Clusters & On-Prem
Senior AI Infrastructure Engineer - GPU Clusters & On-Prem

Spacex • Northern (KY)

Hybrid
USD 165,000 - 265,000
Stock options
401(k) retirement plan
Discretionary bonuses
+7
Senior AI Infra Engineer: GPU & On-Prem Kubernetes
Senior AI Infra Engineer: GPU & On-Prem Kubernetes

Linuxconfig • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
401(k) retirement plan
Medical, vision, dental coverage
+1
Senior AI Infra Engineer: GPU, On-Prem & Kubernetes
Senior AI Infra Engineer: GPU, On-Prem & Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Medical benefits
401(k) plan
Paid parental leave
+1