Senior AI Infrastructure Engineer - GPU Clusters & On-Prem

Spacex

Northern (KY)

Hybrid

USD 165,000 - 265,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) retirement plan
Discretionary bonuses
Health insurance (medical, vision &amp
Dental coverage
Paid vacation
Holidays
Parental leave
Employee stock purchase plan
Disability insurance

Job summary

SpaceX is seeking a Sr. Software Engineer focused on AI infrastructure for Starshield. This role designs, operates, and scales GPU-enabled on-premises infrastructure and Kubernetes-based AI clusters to support national security missions.

You will automate deployments, build scalable software products, and collaborate with AI engineers across the organization on high-availability systems. You will mentor junior engineers and help lead the team toward technical excellence, with a strong emphasis

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement
  • Monitoring and alerting supporting systems to have high availability
  • Identify areas for improvement and create innovative solutions that enable high system availability
  • Mentor and train junior engineers
  • As a senior engineer you must lead the team to technical excellence – your decisions guide the team

Skills

Linux
Kubernetes
Python
Go
C++
Bash

Education

Bachelor’s degree in computer science, information systems/IT, or an engineering discipline

Tools

Terraform
Ansible
OCI containers
Docker

Job description

SpaceX is seeking a Sr. Software Engineer focused on AI infrastructure for Starshield. This role designs, operates, and scales GPU-enabled on-premises infrastructure and Kubernetes-based AI clusters to support national security missions.

You will automate deployments, build scalable software products, and collaborate with AI engineers across the organization on high-availability systems. You will mentor junior engineers and help lead the team toward technical excellence, with a strong emphasis

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — GPU & On-Prem Clusters
Senior AI Infrastructure Engineer — GPU & On-Prem Clusters

Xapply • Washington

On-site
USD 165,000 - 265,000
Stock options
Medical coverage
401(k)
+1
Senior AI Infrastructure Engineer – GPU & Kubernetes
Senior AI Infrastructure Engineer – GPU & Kubernetes

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
Paid vacation
Paid holidays
+2
Senior AI Infrastructure Engineer — GPU & Kubernetes
Senior AI Infrastructure Engineer — GPU & Kubernetes

Engg • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term incentives
401(k) retirement plan
+1
Senior AI Infrastructure Engineer — GPU, Kubernetes, On-Prem
Senior AI Infrastructure Engineer — GPU, Kubernetes, On-Prem

Engg • Redmond (WA)

On-site
USD 165,000 - 270,000
Medical, vision, dental coverage
401(k) with company match
Paid vacation and holidays
+1
Senior AI Infra Engineer: GPU Clusters, Kubernetes & Automation
Senior AI Infra Engineer: GPU Clusters, Kubernetes & Automation

SpaceX • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Senior AI Infra Engineer: GPU-Driven On-Prem AI Clusters
Senior AI Infra Engineer: GPU-Driven On-Prem AI Clusters

InvestedintheMission • Washington

On-site
USD 165,000 - 265,000
Stock options
Comprehensive medical/vision/dental
Paid time off and holidays
Senior AI Infra Engineer — GPU, Kubernetes, On-Prem
Senior AI Infra Engineer — GPU, Kubernetes, On-Prem

SpaceX • Washington

On-site
USD 165,000 - 265,000
Stock awards
401(k)
Health insurance
+3
AI Infrastructure Engineer — GPU & On‑Prem Systems
AI Infrastructure Engineer — GPU & On‑Prem Systems

InvestedintheMission • Palo Alto (CA)

On-site
USD 125,000 - 195,000
Stock options
Health benefits
401(k)
+1
Senior AI Infra Engineer - GPU & Kubernetes
Senior AI Infra Engineer - GPU & Kubernetes

SpaceX • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Senior AI Infra Engineer: GPU, On-Prem & Kubernetes
Senior AI Infra Engineer: GPU, On-Prem & Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Medical benefits
401(k) plan
Paid parental leave
+1