Site Reliability Engineer — AI GPU Infrastructure

United States Digital Space LLC

United States

Remote

USD 125,000 - 195,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
Medical, vision & dental benefits
401(k) retirement plan
Paid vacation & holidays
Parental leave

Job summary

United States Digital Space LLC is seeking a Site Reliability Engineer (STARSHIELD) to design, operate, and scale GPU/CPU infrastructure for Top Secret datacenters and AI clusters. The role focuses on automation, on-prem Kubernetes, and collaboration with AI teams to deliver scalable, reliable software products.

The position requires SRE/DevOps experience, Linux proficiency, and familiarity with Terraform/Ansible. A TS/SCI-like clearance and willingness to travel are important for success.

Qualifications

  • Bachelor's degree in computer science, information systems/IT, or an engineering discipline.
  • 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years in lieu of a degree
  • 1+ years of professional experience with Linux operating systems
  • Experience with Terraform, Ansible, or other infrastructure tools
  • Experience with containerization technologies (i.e. OCI containers, Kubernetes)
  • Experience scripting in Bash, Python, or other similar languages
  • Development experience in Python, C++, or Go

Responsibilities

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms
  • Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
  • Develop automation to deploy and manage on-premise Kubernetes/AI clusters, and operating systems
  • Deploy and manage core infrastructure such as databases, monitoring and distributed storage
  • Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
  • Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement

Skills

Linux
Terraform
Ansible
Kubernetes
Python
Go
C++

Education

Bachelor's degree in computer science, information systems/IT, or engineering

Tools

OCI containers
Makefiles
Bazel

Job description

United States Digital Space LLC is seeking a Site Reliability Engineer (STARSHIELD) to design, operate, and scale GPU/CPU infrastructure for Top Secret datacenters and AI clusters. The role focuses on automation, on-prem Kubernetes, and collaboration with AI teams to deliver scalable, reliable software products.

The position requires SRE/DevOps experience, Linux proficiency, and familiarity with Terraform/Ansible. A TS/SCI-like clearance and willingness to travel are important for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI GPU Infra & On-Prem Systems
Senior SRE: AI GPU Infra & On-Prem Systems

United States Digital Space LLC • United States

Remote
USD 165,000 - 265,000
Stock options / long-term incentives
Health, vision & dental coverage
401(k) retirement plan
+3
SRE: AI GPU Infra for High-Security Clusters
SRE: AI GPU Infra for High-Security Clusters

United States Digital Space LLC • Washington, El Segundo (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k) plan
Medical, vision, dental coverage
+1
Site Reliability Engineer - AI GPU Infrastructure
Site Reliability Engineer - AI GPU Infrastructure

Spacex • Washington

On-site
USD 125,000 - 195,000
Stock options
Long-term incentives
Health, vision and dental coverage
+3
Senior SRE for AI GPU Infra (Top Secret)
Senior SRE for AI GPU Infra (Top Secret)

SpaceX • Redmond (WA)

On-site
USD 165,000 - 270,000
Stock options and long-term incentives
Comprehensive medical/dental/vision
SRE: AI GPU Infra & On-Prem Kubernetes
SRE: AI GPU Infra & On-Prem Kubernetes

SpaceX • Washington

On-site
USD 125,000 - 195,000
Stock options
Medical, vision, dental
401(k) retirement plan
+2
AI Infrastructure SRE — GPU & On-Prem Kubernetes
AI Infrastructure SRE — GPU & On-Prem Kubernetes

SPACE EXPLORATION TECHNOLOGIES CORP • Washington

On-site
USD 125,000 - 195,000
Stock options/long-term incentives
Medical, Vision, Dental
401(k) plan
+2
SRE: AI GPU Infrastructure for Starshield
SRE: AI GPU Infrastructure for Starshield

SpaceX • Redmond (WA)

On-site
USD 125,000 - 200,000
Medical insurance
Vision insurance
Dental coverage
+6
Site Reliability Engineer — AI Infra & GPU Clusters
Site Reliability Engineer — AI Infra & GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 195,000
Stock options
401(k)
Health coverage
+1
AI Infrastructure SRE - GPU Clusters
AI Infrastructure SRE - GPU Clusters

SPACE EXPLORATION TECHNOLOGIES CORP • Redmond (WA), Northern (KY)

Hybrid
USD 125,000 - 200,000
Stock options
401(k) plan
Comprehensive medical benefits
+1
SRE – AI Infrastructure & 100k+ GPU Clusters
SRE – AI Infrastructure & 100k+ GPU Clusters

SpaceX • Union Hill-Novelty Hill (WA)

On-site
USD 125,000 - 200,000
Stock options and long-term incentives
401(k) retirement plan
Health, vision, dental coverage
+4