Senior AI Runtime Engineer — Scalable GPU Training

Databricks

California (MO)

On-site

USD 160,000 - 225,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Databricks is hiring a Senior Software Engineer for AI Runtime to architect and scale the GPU training stack. You will drive multi-node orchestration, fault tolerance, and performance optimization for large-scale AI training workloads across GPUs and clusters.

You will mentor engineers, collaborate with product and research teams, and help shape the APIs and tooling that enable customers to launch and monitor production training jobs.

Qualifications

  • 5+ years of experience building and operating large-scale distributed systems.
  • Experience with distributed training frameworks (PyTorch, FSDP, DeepSpeed, Megatron).
  • Strong understanding of training resilience patterns (checkpointing, failure detection, automatic recovery).
  • Solid grasp of GPU performance fundamentals (accelerator architecture, NVLink/InfiniBand).
  • Experience building and operating managed, multi-tenant cloud platform products.

Responsibilities

  • Architect and evolve AIR's managed GPU training platform for scale and resilience.
  • Mentor engineers and collaborate across product, research, and platform teams.
  • Shape APIs, CLI, and developer experience for launching and monitoring training jobs.
  • Lead end-to-end engineering efforts from design to production rollout.
  • Contribute to Databricks' AI training infrastructure direction.

Skills

Distributed systems
GPU training infra
Data parallelism
System design
Communication
Mentoring

Education

BS in Computer Science
MS or PhD preferred

Tools

PyTorch
FSDP
DeepSpeed
Megatron

Job description

Databricks is hiring a Senior Software Engineer for AI Runtime to architect and scale the GPU training stack. You will drive multi-node orchestration, fault tolerance, and performance optimization for large-scale AI training workloads across GPUs and clusters.

You will mentor engineers, collaborate with product and research teams, and help shape the APIs and tooling that enable customers to launch and monitor production training jobs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Runtime Engineer - Scalable GPU Training Platform
Staff AI Runtime Engineer - Scalable GPU Training Platform

Databricks • California (MO)

On-site
USD 190,000 - 265,000
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime Engineer: Scale GPU Training
Senior AI Runtime Engineer: Scale GPU Training

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Platform & GPU Training Manager
Senior AI Platform & GPU Training Manager

Databricks • California (MO)

On-site
USD 229,000 - 297,000