Senior AI Runtime Platform Leader

jobr.pro

Mountain View (CA)

On-site

USD 228,600 - 297,120

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Databricks is seeking a Senior Engineering Manager to lead the AIR product team—owning both user-facing experiences and the underlying GPU training infrastructure. You will shape the roadmap, drive scalability, and coordinate across platform, product, infrastructure, and research groups to deliver high-impact, long-running training workloads.

You will mentor a high-performing team, emphasize reliability, observability, and performance, and partner with recruiting to attract top engineering

Qualifications

  • 8+ years of software engineering experience, with 3+ years in engineering management.
  • Track record building and operating managed GPU training infrastructure at scale (hundreds to thousands of GPUs).
  • Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies.
  • Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long-running jobs.
  • Understanding of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
  • Experience building platform products with clear SLAs and an owned customer experience.
  • Strong cross-functional leadership and excellent collaboration across engineering, product, and research teams.

Responsibilities

  • Lead, mentor, and grow a high-performing engineering team responsible for the Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate closely with product, research, platform, infrastructure teams, and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
  • Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
  • Partner with recruiting to attract, hire, and develop top-tier engineering talent.

Skills

GPU training infra
Distributed training frameworks
Platform product ownership
Cross-functional leadership
Performance optimization
Reliability & observability
Team recruiting collaboration

Education

BS/MS in CS/EE or related field

Job description

Databricks is seeking a Senior Engineering Manager to lead the AIR product team—owning both user-facing experiences and the underlying GPU training infrastructure. You will shape the roadmap, drive scalability, and coordinate across platform, product, infrastructure, and research groups to deliver high-impact, long-running training workloads.

You will mentor a high-performing team, emphasize reliability, observability, and performance, and partner with recruiting to attract top engineering

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (WY)

On-site
USD 229,000 - 297,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
AI Runtime Engineering Lead for Scalable GPU Training
AI Runtime Engineering Lead for Scalable GPU Training

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

United States Digital Space LLC • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Engineering Manager, AI-Driven CX Platform
Senior Engineering Manager, AI-Driven CX Platform

Neura Market • Mountain View (CA)

Hybrid
USD 222,000 - 313,000
Sr. Engineering Manager, AI Runtime
Sr. Engineering Manager, AI Runtime

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000