AI Runtime Engineering Lead for Scalable GPU Training

United States Digital Space LLC

Mountain View (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking a Senior Engineering Manager to lead the team responsible for the AIR product's training infrastructure and customer-facing experience. You will shape the architecture for scalable GPU training, drive the roadmap, and collaborate with product, research, platform, and customers to deliver end-to-end solutions.

You will mentor a high-performing team, implement reliability and observability practices for long-running multi-node jobs, and partner with

Qualifications

  • 8+ years of software engineering experience, with 3+ years in engineering management.
  • Experience building and operating managed GPU training infrastructure at scale (hundreds to thousands of GPUs).
  • Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline).
  • Experience with training resilience: checkpointing, elastic training, failure recovery for long-running jobs.
  • Understanding of GPU performance fundamentals (NCCL, interconnects, memory optimization).

Responsibilities

  • Lead, mentor, and grow a high-performing engineering team responsible for the Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate with product, research, platform, infrastructure teams, and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
  • Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
  • Partner with recruiting to attract, hire, and develop top-tier engineering talent.

Skills

Leadership
Distributed training
PyTorch
DeepSpeed
Megatron-LM
NCCL
GPU performance

Education

BS/MS in CS/EE

Tools

Cluster orchestration
NVIDIA GPUs
Tensor parallelism

Job description

United States Digital Space LLC is seeking a Senior Engineering Manager to lead the team responsible for the AIR product's training infrastructure and customer-facing experience. You will shape the architecture for scalable GPU training, drive the roadmap, and collaborate with product, research, platform, and customers to deliver end-to-end solutions.

You will mentor a high-performing team, implement reliability and observability practices for long-running multi-node jobs, and partner with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (WY)

On-site
USD 229,000 - 297,000
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Staff Engineer, GPU Inference & Training Platform
Staff Engineer, GPU Inference & Training Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
AI Solutions Architecture Lead
AI Solutions Architecture Lead

Socket.dev • Santa Clara (UT)

On-site
USD 224,000 - 357,000