GPU Infrastructure Engineer - Scale & Automation

United States Digital Space LLC

United States

Remote

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking a Software Engineer to design, build, and operate software for large-scale GPU infrastructure powering ChatGPT inference. You will collaborate with infrastructure, research, and product teams to improve reliability and compute utilization.

The role emphasizes automating fleet operations, improving observability, and delivering scalable solutions across thousands of GPUs in a fast-moving environment.

Qualifications

  • 5+ years of software engineering experience building production infrastructure.
  • Strong programming skills in Go, Python, C++, Rust, or similar systems languages.
  • Experience designing and operating highly available distributed systems.
  • Experience with GPU infrastructure, high‑performance computing, ML infrastructure, or large‑scale compute platforms.
  • Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
  • Excellent debugging, systems design, and operational problem‑solving skills.
  • Strong communication skills and experience collaborating across engineering organizations.

Responsibilities

  • Design, build, and operate software that manages large‑scale GPU infrastructure supporting ChatGPT inference.
  • Build internal platforms, tooling, and AI‑powered agents that automate fleet operations and reduce operational overhead.
  • Improve observability, reliability, and operational efficiency across thousands of GPUs.
  • Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
  • Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.
  • Partner with research, platform, networking, and systems teams to improve our compute platform.
  • Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.

Skills

Distributed systems
Software engineering
Cross-team collaboration
Problem solving
System design

Tools

Go
Python
C++
Rust
Kubernetes
Linux
Cloud infrastructure
Observability tooling

Job description

United States Digital Space LLC is seeking a Software Engineer to design, build, and operate software for large-scale GPU infrastructure powering ChatGPT inference. You will collaborate with infrastructure, research, and product teams to improve reliability and compute utilization.

The role emphasizes automating fleet operations, improving observability, and delivering scalable solutions across thousands of GPUs in a fast-moving environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infra Engineer for Scalable AI Compute
GPU Infra Engineer for Scalable AI Compute

OpenAI • United States

On-site
USD 180,000 - 240,000
Staff Engineer, GPU Inference & Training Platform
Staff Engineer, GPU Inference & Training Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior AI Infrastructure Engineer | Scale GPU Clusters
Senior AI Infrastructure Engineer | Scale GPU Clusters

Fuel Talent LLC • Seattle (WA)

Hybrid
USD 126,000 - 189,000
GPU-Driven AI Infrastructure & Platform Engineer
GPU-Driven AI Infrastructure & Platform Engineer

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Competitive compensation package
Professional development and training
Conferences and working groups
+1
Senior System Engineer - AI Inference & GPU Performance
Senior System Engineer - AI Inference & GPU Performance

United States Digital Space LLC • United States

Remote
USD 150,000 - 230,000
Competitive compensation
Career growth and learningOpportun it
Remote-Ready Compute Engineer for Scalable GPU Fleets
Remote-Ready Compute Engineer for Scalable GPU Fleets

Insight Global • Town of Texas (WI)

On-site
USD 120,000 - 180,000
GPU Cluster Infra Engineer - Reliability & Automation
GPU Cluster Infra Engineer - Reliability & Automation

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
Software Engineer, GPU Infrastructure- ChatGPT Engineering
Software Engineer, GPU Infrastructure- ChatGPT Engineering

OpenAI • United States

On-site
USD 180,000 - 240,000
GPU Systems Engineer — Distributed Training & Inference
GPU Systems Engineer — Distributed Training & Inference

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays