Staff AI Runtime Architect - Scale GPU Training

Databricks

Mountain View (CA)

On-site

USD 190,000 - 265,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual performance bonus
Equity options
Comprehensive benefits package

Job summary

Databricks is looking for a Staff Software Engineer for AI Runtime in Mountain View, California, to help build and scale the systems necessary for large-scale GPU training. You will drive the architecture of a managed GPU training platform ensuring high throughput and resilience across many accelerators.

The ideal candidate will have over 10 years of experience in distributed systems and machine learning, with a solid grasp of GPU performance. Join Databricks to make a significant impact in the AI training domain.

Qualifications

  • 10+ years of experience in large-scale distributed systems.
  • Hands-on experience with distributed training frameworks.
  • Understanding of GPU performance fundamentals.

Responsibilities

  • Drive the architecture of AIR’s managed GPU training platform.
  • Solve large-scale training problems.
  • Lead engineering efforts from design to production.

Skills

Large-scale distributed systems
GPU training infrastructure
Machine Learning systems
Distributed training frameworks
High-performance computing

Education

BS in Computer Science or related field
MS or PhD preferred

Tools

PyTorch
FSDP
DeepSpeed
Megatron

Job description

Databricks is looking for a Staff Software Engineer for AI Runtime in Mountain View, California, to help build and scale the systems necessary for large-scale GPU training. You will drive the architecture of a managed GPU training platform ensuring high throughput and resilience across many accelerators.

The ideal candidate will have over 10 years of experience in distributed systems and machine learning, with a solid grasp of GPU performance. Join Databricks to make a significant impact in the AI training domain.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime Engineer: Scale GPU Training
Senior AI Runtime Engineer: Scale GPU Training

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (WY)

On-site
USD 229,000 - 297,000
Senior GenAI Kernel & GPU Optimization Engineer
Senior GenAI Kernel & GPU Optimization Engineer

Databricks • Mountain View (CA)

On-site
USD 166,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Staff AI Research Infrastructure Engineer
Staff AI Research Infrastructure Engineer

Doist • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 270,000