Senior Engineering Manager, AI Runtime & GPU Training

Databricks

Mountain View (CA)

On-site

USD 228,600 - 297,120

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Databricks is seeking a Senior Engineering Manager to lead the AIR team responsible for the product experience and foundational infrastructure of GPU training at scale. You will shape the roadmap, drive end-to-end delivery, and collaborate across platform, product, research, and customers to ensure reliable, high-performance training pipelines.

You will mentor a high-performing team, influence architecture, and implement observability and resilience practices for long-running multi-node jobs,

Qualifications

  • 8+ years of software engineering and 3+ years in management.
  • Experience building and operating GPU training infrastructure at scale (hundreds/thousands of GPUs).
  • Deep familiarity with distributed training frameworks: PyTorch, DeepSpeed, Megatron LM; parallelism such as FSDP.
  • Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long running jobs.
  • Knowledge of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
  • Experience building platform products with clear SLAs and owning the customer experience.
  • Strong cross-functional leadership across platform, product, and research, with ability to lead through ambiguity.
  • Excellent collaboration and communication across engineering, product, and research organizations.
  • BS/MS in Computer Science, Electrical Engineering, or related technical field.

Responsibilities

  • Lead, mentor, and grow a high performing engineering team responsible for the Custom Training product and its foundational infrastructure.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate with product, research, platform, infrastructure teams, and customers to drive end to end delivery from ideation to launch and operation.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
  • Build observability and reliability practices for long running, multi node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
  • Partner with recruiting to attract, hire, and develop top tier engineering talent.

Skills

Software engineering
Engineering management
GPU training infra
PyTorch
DeepSpeed
Megatron LM
FSDP
Training resilience
NCCL & memory
SLAs & customer focus
Cross-functional leadership
Communication

Education

BS/MS in Computer Science or Electrical Engineering

Job description

Databricks is seeking a Senior Engineering Manager to lead the AIR team responsible for the product experience and foundational infrastructure of GPU training at scale. You will shape the roadmap, drive end-to-end delivery, and collaborate across platform, product, research, and customers to ensure reliable, high-performance training pipelines.

You will mentor a high-performing team, influence architecture, and implement observability and resilience practices for long-running multi-node jobs,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,600 - 297,120
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,600 - 297,120
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

jobr.pro • Mountain View (CA)

On-site
USD 228,600 - 297,120
Senior AI Runtime Engineer: Scale GPU Training
Senior AI Runtime Engineer: Scale GPU Training

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

Menlo Ventures • San Francisco (CA)

On-site
USD 228,600 - 297,120
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

Databricks • Mountain View (CA)

On-site
USD 228,600 - 297,120
Senior Engineering Manager, AI/BI Dashboard Platform
Senior Engineering Manager, AI/BI Dashboard Platform

Xapply • Mountain View (CA), Northern (KY)

Hybrid
USD 229,000 - 314,000