Senior AI Infra Engineer: GPU & On-Prem Kubernetes

Linuxconfig

Redmond (WA)

On-site

USD 165,000 - 270,000

Full time

30 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) retirement plan
Medical, vision, dental coverage
Paid time off and holidays

Job summary

SpaceX is seeking a Sr. Software Engineer, AI Infrastructure (Starshield) in Redmond, WA to design, operate, and scale GPU-backed AI infrastructure supporting critical national security missions.

The role encompasses building automated on-premise Kubernetes clusters, managing GPU services, and collaborating with AI engineers to deliver highly scalable software products. Expect to mentor junior engineers and drive technical excellence.

Qualifications

  • Bachelor’s degree in computer science, information systems/IT, or an engineering discipline and 5+ years of professional experience with Linux operating systems; OR 7+ years in lieu of a degree
  • 5+ year of experience with Kubernetes
  • 5+ year of experience managing Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go
  • Preferred: 5+ years of experience with Python and Python-based development frameworks
  • Knowledge of Linux boot process and systems configuration
  • Testing, CI, build, deployment & monitoring familiarity
  • Build technologies like Bazel and Makefiles
  • Experience automating thousands of servers (Terraform/Ansible)
  • Strong TCP/IP networking knowledge
  • Cloud virtualization experience
  • Excellent communications skills
  • Active security clearance requirements

Responsibilities

  • Manage and provide support for GPU as a service on bare metal hardware and virtualized platforms
  • Design, validate, and productize AI cluster solutions at large scale (100k+ GPUs)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters and OSes
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Collaborate with AI engineers to deliver scalable, operable products
  • Oversee lifecycle of services from design to deployment and refinement
  • Implement monitoring and alerting for high availability
  • Identify improvement opportunities and implement innovative solutions
  • Mentor and train junior engineers
  • Lead the team to technical excellence and strategic direction

Skills

Kubernetes
Linux
Terraform
Ansible
OCI containers
Bash
Python
C++
Go
GPGPU/AI infra

Education

Bachelor’s degree in CS/IT/Engineering

Tools

Terraform
Ansible
Kubernetes
OCI containers

Job description

SpaceX is seeking a Sr. Software Engineer, AI Infrastructure (Starshield) in Redmond, WA to design, operate, and scale GPU-backed AI infrastructure supporting critical national security missions.

The role encompasses building automated on-premise Kubernetes clusters, managing GPU services, and collaborating with AI engineers to deliver highly scalable software products. Expect to mentor junior engineers and drive technical excellence.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — GPU, Kubernetes, On-Prem
Senior AI Infrastructure Engineer — GPU, Kubernetes, On-Prem

Engg • Redmond (WA)

On-site
USD 165,000 - 270,000
Medical, vision, dental coverage
401(k) with company match
Paid vacation and holidays
+1
AI Infrastructure Engineer — GPU & On-Prem Kubernetes
AI Infrastructure Engineer — GPU & On-Prem Kubernetes

InvestedintheMission • Redmond (WA)

On-site
USD 125,000 - 200,000
Company stock
401(k) plan
Paid time off
AI Infrastructure Engineer — GPU & On-Prem Kubernetes
AI Infrastructure Engineer — GPU & On-Prem Kubernetes

Spacex • Redmond (WA)

On-site
USD 125,000 - 200,000
Long‑term incentives
Stock options
Bonuses
+5
Senior AI Infra Engineer: GPU, On-Prem & Kubernetes
Senior AI Infra Engineer: GPU, On-Prem & Kubernetes

InvestedintheMission • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Medical benefits
401(k) plan
Paid parental leave
+1
Senior AI Infrastructure Engineer – GPU & Kubernetes
Senior AI Infrastructure Engineer – GPU & Kubernetes

InvestedintheMission • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
Paid vacation
Paid holidays
+2
Senior AI Infrastructure Engineer (GPU/Kubernetes)
Senior AI Infrastructure Engineer (GPU/Kubernetes)

Spacex • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options
Medical/vision/dental
401(k) retirement plan
+3
Senior AI Infra Engineer - GPU & Kubernetes
Senior AI Infra Engineer - GPU & Kubernetes

SpaceX • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Senior AI Infra Engineer: GPU & Kubernetes Scale Leader
Senior AI Infra Engineer: GPU & Kubernetes Scale Leader

Engg • Palo Alto (CA)

On-site
USD 165,000 - 265,000
Health insurance
Dental coverage
Vision coverage
+7
AI Infrastructure Engineer — GPU & On-Prem Kubernetes
AI Infrastructure Engineer — GPU & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 190,000
Stock options and long‑term incentives
Discretionary bonuses
Comprehensive medical, vision, dental
Senior AI Infrastructure Engineer — GPU & Kubernetes
Senior AI Infrastructure Engineer — GPU & Kubernetes

Engg • Hawthorne (CA)

On-site
USD 165,000 - 265,000
Stock options
Long-term incentives
401(k) retirement plan
+1