Staff AI Runtime Engineer - Scalable GPU Training Platform

Databricks

California (MO)

On-site

USD 190,000 - 265,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Databricks is seeking a Staff Software Engineer for AI Runtime to scale large‑scale GPU training systems. You will architect the AIR training stack, drive multi-node orchestration, and improve resilience, performance, and developer experience.

You will mentor engineers and align technical work with a long‑term vision across product, research, and platform teams. You will collaborate with diverse teams to enable production training at scale, support accelerator readiness, and expand capabilities

Qualifications

  • 10+ years of experience building and operating large-scale distributed systems, with depth in GPU infrastructure or ML systems.
  • Hands-on experience with distributed training frameworks and parallelism strategies for large models.
  • Strong understanding of training resilience, including checkpointing, failure detection, and automatic recovery.

Responsibilities

  • Lead architecture and evolution of AIR's managed GPU training stack to deliver scalable, high-throughput training across fleets.
  • Drive multi-node orchestration, distributed parallelism, GPU scheduling, data loading, and long-running job resilience.
  • Push GPU efficiency and training performance while reducing cost per training run across architectures and hardware generations.
  • Build resilience and observability foundations to keep multi-node jobs healthy and recoverable from failures.

Skills

Distributed training
GPU training infrastructure
PyTorch/DeepSpeed/Megatron
Distributed systems design
Cloud multi-tenant platforms
Performance optimization
Architectural leadership
Mentoring engineers
Strong communication
BS in CS (MS/PhD preferred)

Education

BS in Computer Science or related field

Job description

Databricks is seeking a Staff Software Engineer for AI Runtime to scale large‑scale GPU training systems. You will architect the AIR training stack, drive multi-node orchestration, and improve resilience, performance, and developer experience.

You will mentor engineers and align technical work with a long‑term vision across product, research, and platform teams. You will collaborate with diverse teams to enable production training at scale, support accelerator readiness, and expand capabilities

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Runtime Engineer — Scalable GPU Training
Senior AI Runtime Engineer — Scalable GPU Training

Databricks • California (MO)

On-site
USD 160,000 - 225,000
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Platform & GPU Training Manager
Senior AI Platform & GPU Training Manager

Databricks • California (MO)

On-site
USD 229,000 - 297,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000
Senior AI Runtime Engineer: Scale GPU Training
Senior AI Runtime Engineer: Scale GPU Training

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package