Senior Engineering Manager, AI Runtime & GPU Training

Databricks

Mountain View (CA)

On-site

USD 228,600 - 297,120

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Databricks is seeking a Senior Engineering Manager to lead the AIR team responsible for the product experience and foundational infrastructure of GPU training at scale. You will shape the roadmap, drive end-to-end delivery, and collaborate across platform, product, research, and customers to ensure reliable, high-performance training pipelines.

You will mentor a high-performing team, influence architecture, and implement observability and resilience practices for long-running multi-node jobs,

Qualifications

  • 8+ years of software engineering and 3+ years in management.
  • Experience building and operating GPU training infrastructure at scale (hundreds/thousands of GPUs).
  • Deep familiarity with distributed training frameworks: PyTorch, DeepSpeed, Megatron LM; parallelism such as FSDP.
  • Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long running jobs.
  • Knowledge of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
  • Experience building platform products with clear SLAs and owning the customer experience.
  • Strong cross-functional leadership across platform, product, and research, with ability to lead through ambiguity.
  • Excellent collaboration and communication across engineering, product, and research organizations.
  • BS/MS in Computer Science, Electrical Engineering, or related technical field.

Responsibilities

  • Lead, mentor, and grow a high performing engineering team responsible for the Custom Training product and its foundational infrastructure.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate with product, research, platform, infrastructure teams, and customers to drive end to end delivery from ideation to launch and operation.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
  • Build observability and reliability practices for long running, multi node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
  • Partner with recruiting to attract, hire, and develop top tier engineering talent.

Skills

Software engineering
Engineering management
GPU training infra
PyTorch
DeepSpeed
Megatron LM
FSDP
Training resilience
NCCL & memory
SLAs & customer focus
Cross-functional leadership
Communication

Education

BS/MS in Computer Science or Electrical Engineering

Job description

Databricks is seeking a Senior Engineering Manager to lead the AIR team responsible for the product experience and foundational infrastructure of GPU training at scale. You will shape the roadmap, drive end-to-end delivery, and collaborate across platform, product, research, and customers to ensure reliable, high-performance training pipelines.

You will mentor a high-performing team, influence architecture, and implement observability and resilience practices for long-running multi-node jobs,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (WY)

On-site
USD 229,000 - 297,000
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
AI Runtime Engineering Lead for Scalable GPU Training
AI Runtime Engineering Lead for Scalable GPU Training

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000