Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures

San Francisco (CA)

On-site

USD 228,600 - 297,120

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Databricks is seeking a Senior Engineering Manager to lead the team owning the AIR product experience and its foundational GPU training infrastructure. You will shape customer‑facing capabilities while ensuring scalability, performance, and reliability across multi‑node training at scale.

You will collaborate with platform, product, infrastructure, and research teams to drive end‑to‑end delivery, mentor engineers, and own architectural decisions for state‑of‑the art GPU training solutions with

Qualifications

  • 8+ years of software engineering experience with 3+ years in engineering management.
  • Track record building and operating managed GPU training infrastructure at scale (100s/1000s GPUs).
  • Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism).
  • Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long-running jobs.
  • Understanding of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
  • Experience building platform products with clear SLAs where you've owned the customer experience, not just the backend.
  • Strong cross-functional leadership across platform, product, and research teams, with the ability to lead through ambiguity and deliver complex projects.
  • Excellent collaboration and communication skills across engineering, product, and research organizations.
  • BS/MS in Computer Science, Electrical Engineering, or related technical field.

Responsibilities

  • Lead, mentor, and grow a high‑performing engineering team responsible for the Custom Training product and its foundational infrastructure.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate with product, research, platform, and infrastructure teams, and customers to drive end‑to‑end delivery from ideation to launch.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, translating engineering decisions into product impact.
  • Build observability and reliability practices for long‑running, multi‑node training jobs, including checkpoint strategies and failure recovery.
  • Partner with recruiting to attract, hire, and develop top‑tier engineering talent.

Skills

Engineering management
Distributed training infra
PyTorch
DeepSpeed
Megatron-LM
NCCL
GPU performance tuning
Cross-functional leadership
Platform products

Education

BS/MS in Computer Science or related field

Tools

PyTorch
DeepSpeed
Megatron-LM

Job description

Databricks is seeking a Senior Engineering Manager to lead the team owning the AIR product experience and its foundational GPU training infrastructure. You will shape customer‑facing capabilities while ensuring scalability, performance, and reliability across multi‑node training at scale.

You will collaborate with platform, product, infrastructure, and research teams to drive end‑to‑end delivery, mentor engineers, and own architectural decisions for state‑of‑the art GPU training solutions with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (WY)

On-site
USD 229,000 - 297,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
AI Runtime Engineering Lead for Scalable GPU Training
AI Runtime Engineering Lead for Scalable GPU Training

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior AI Runtime Engineer: Scale GPU Training
Senior AI Runtime Engineer: Scale GPU Training

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000