AI Infrastructure Engineer - GPU & On-Prem

InvestedintheMission

Washington (District of Columbia)

On-site

USD 125,000 - 195,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
Bonuses
Medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
3 weeks vacation + 10+ holidays

Job summary

SpaceX is hiring a Software Engineer to design, operate and scale Starshield AI infrastructure. You will manage GPU/CPU infrastructure deployments in Top Secret datacenters and build automation for on-prem Kubernetes/AI clusters.

You will collaborate with AI engineers to deliver scalable products, maintain critical services, and ensure high availability. This role requires readiness to obtain and maintain a Top Secret clearance and travel as needed.

Qualifications

  • 1+ years of professional experience in site reliability engineering or DevOps
  • 1+ years of professional experience with Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions that enable high system availability

Skills

Linux
Terraform
Ansible
Kubernetes
Bash
Python
C++
Go

Education

Bachelor's degree in CS/IT/Engineering

Tools

OCI containers
Docker
Kubernetes (cluster management)

Job description

SpaceX is hiring a Software Engineer to design, operate and scale Starshield AI infrastructure. You will manage GPU/CPU infrastructure deployments in Top Secret datacenters and build automation for on-prem Kubernetes/AI clusters.

You will collaborate with AI engineers to deliver scalable products, maintain critical services, and ensure high availability. This role requires readiness to obtain and maintain a Top Secret clearance and travel as needed.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: GPU, On‑Prem & Kubernetes
AI Infrastructure Engineer: GPU, On‑Prem & Kubernetes

Engg • Redmond (WA)

On-site
USD 125,000 - 200,000
Stock options
Medical, vision & dental coverage
401(k) plan
+2
AI Infrastructure Engineer — GPU & On‑Prem Systems
AI Infrastructure Engineer — GPU & On‑Prem Systems

InvestedintheMission • Palo Alto (CA)

On-site
USD 125,000 - 195,000
Stock options
Health benefits
401(k)
+1
AI Infrastructure Engineer – GPU & On-Prem
AI Infrastructure Engineer – GPU & On-Prem

InvestedintheMission • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Senior AI Infra Engineer: GPU-Driven On-Prem AI Clusters
Senior AI Infra Engineer: GPU-Driven On-Prem AI Clusters

InvestedintheMission • Washington

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical/vision/dental
Paid time off and holidays
AI Infrastructure Engineer — GPU & On-Prem Kubernetes
AI Infrastructure Engineer — GPU & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 190,000
Stock options and long‑term incentives
Discretionary bonuses
Comprehensive medical, vision, dental
Senior AI Infrastructure Engineer - GPU Clusters & On-Prem
Senior AI Infrastructure Engineer - GPU Clusters & On-Prem

Spacex • Northern (KY)

Hybrid
USD 165,000 - 265,000
Stock options
401(k) retirement plan
Discretionary bonuses
+7
AI Infrastructure Engineer – GPU & On-Prem Cloud
AI Infrastructure Engineer – GPU & On-Prem Cloud

Linuxconfig • Hawthorne (CA), Northern (KY)

Hybrid
USD 125,000 - 195,000
Stock options
Medical, vision & dental coverage
401(k) retirement plan
+1
AI Infrastructure Engineer — GPU Clusters & SRE
AI Infrastructure Engineer — GPU Clusters & SRE

SpaceX • Palo Alto (CA)

On-site
USD 125,000 - 195,000
Senior AI Infrastructure Engineer — GPU & On-Prem Clusters
Senior AI Infrastructure Engineer — GPU & On-Prem Clusters

Xapply • Washington

On-site
USD 165,000 - 265,000
Stock options
Medical coverage
401(k)
+1
AI Infra Engineer: Scale GPU Clusters & On-Prem AI
AI Infra Engineer: Scale GPU Clusters & On-Prem AI

Xapply • Palo Alto (CA)

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Medical, vision & dental
+5