Staff Engineer, GPU Inference & Training Platform

United States Digital Space LLC

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking an infrastructure leader to own a self-serve GPU compute platform for training and inference workloads. You will design and operate the system that lets researchers launch jobs across multi-cloud GPU fleets without manual provisioning.

You will own provisioning, scheduling, and reliability across providers, building fault-tolerance, observability, and a coherent platform roadmap for scalable AI workloads.

Qualifications

  • Deep Kubernetes experience with custom operators and CRDs.
  • Experience managing GPU clusters at scale with NVIDIA GPUs and high-speed networks.
  • Experience orchestrating compute across multiple clouds.
  • Strong distributed systems fundamentals.
  • Proficiency in Go, Rust or C++ for infrastructure.
  • Experience with long-running training and high-availability inference.
  • Ability to own problems end-to-end.

Responsibilities

  • Build a self-serve compute platform for training and inference workloads.
  • Operate the GPU fleet across providers with provisioning and lifecycle management.
  • Develop scheduling and placement to balance capacity across clouds.
  • Support both long-running training jobs and production inference workloads.
  • Own Kubernetes orchestration across multiple clusters and providers.
  • Implement fault tolerance, autoscaling, and observability.
  • Set technical direction across teams to align platform roadmap.

Skills

Kubernetes
GPU clusters
Multi-cloud orchestration
Distributed systems
Go/Rust/C++

Tools

Custom Kubernetes Operators

Job description

United States Digital Space LLC is seeking an infrastructure leader to own a self-serve GPU compute platform for training and inference workloads. You will design and operate the system that lets researchers launch jobs across multi-cloud GPU fleets without manual provisioning.

You will own provisioning, scheduling, and reliability across providers, building fault-tolerance, observability, and a coherent platform roadmap for scalable AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU-Driven AI Infrastructure & Platform Engineer
GPU-Driven AI Infrastructure & Platform Engineer

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Competitive compensation package
Professional development and training
Conferences and working groups
+1
GPU Infrastructure Engineer - Scale & Automation
GPU Infrastructure Engineer - Scale & Automation

United States Digital Space LLC • United States

Remote
USD 150,000 - 210,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior System Engineer - AI Inference & GPU Performance
Senior System Engineer - AI Inference & GPU Performance

United States Digital Space LLC • United States

Remote
USD 150,000 - 230,000
Competitive compensation
Career growth and learningOpportun it
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Platform Engineer — GPU Inference & Training
Senior Platform Engineer — GPU Inference & Training

Neura Market • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior Platform Engineer - GPU Inference & Training
Senior Platform Engineer - GPU Inference & Training

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Lead, Multi-Cloud GPU Inference & Training
Platform Lead, Multi-Cloud GPU Inference & Training

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000