Senior AI Platform & GPU Training Manager

Databricks

California (MO)

On-site

USD 229,000 - 297,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Databricks is seeking a Senior Engineering Manager to lead the team owning the AIR product experience and its foundational infrastructure. You will shape customer-facing capabilities while ensuring scalability, extensibility, and performance of GPU training and related areas, collaborating across platform, product, infrastructure, and research.

You will mentor an engineering team, define AIR’s roadmap, drive end-to-end delivery from ideation to launch, and build observability and reliability for

Qualifications

  • 8+ years of software engineering experience with 3+ years in engineering management.
  • Track record building and operating GPU training infrastructure at scale (100s/1000s GPUs).
  • Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism).
  • Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long-running jobs.
  • Understanding of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
  • Experience building platform products with clear SLAs where you've owned the customer experience, not just the backend.
  • Strong cross-functional leadership across platform, product, and research teams, with the ability to lead through ambiguity and deliver complex projects.
  • Excellent collaboration and communication skills across engineering, product, and research organizations.
  • BS/MS in Computer Science, Electrical Engineering, or related technical field.

Responsibilities

  • Lead, mentor, and grow a high-performing engineering team responsible for the Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate closely with product, research, platform, infrastructure teams, and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
  • Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
  • Partner with recruiting to attract, hire, and develop top-tier engineering talent.

Skills

Engineering management
Distributed training
GPU training infra
PyTorch
DeepSpeed
Megatron-LM
FSDP
Checkpointing
NCCL
Cross-functional leadership
Communication

Education

BS/MS in CS/EE

Tools

NCCL

Job description

Databricks is seeking a Senior Engineering Manager to lead the team owning the AIR product experience and its foundational infrastructure. You will shape customer-facing capabilities while ensuring scalability, extensibility, and performance of GPU training and related areas, collaborating across platform, product, infrastructure, and research.

You will mentor an engineering team, define AIR’s roadmap, drive end-to-end delivery from ideation to launch, and build observability and reliability for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Staff AI Runtime Engineer - Scalable GPU Training Platform
Staff AI Runtime Engineer - Scalable GPU Training Platform

Databricks • California (MO)

On-site
USD 190,000 - 265,000
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime Engineer — Scalable GPU Training
Senior AI Runtime Engineer — Scalable GPU Training

Databricks • California (MO)

On-site
USD 160,000 - 225,000
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000