Senior AI Runtime & GPU Training Systems Engineer

Cacheflow

San Francisco (CA)

On-site

USD 190,000 - 265,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Annual performance bonus
Comprehensive benefits

Job summary

Databricks is seeking a Staff Software Engineer to contribute to the AI Runtime platform for large-scale GPU training. The role centers around driving the architecture and scalability of the managed training stack, ensuring fast and reliable operations.

Candidates should have over 10 years of experience with distributed systems and GPU training, along with a passion for mentoring and engineering excellence. This position is based in San Francisco, California.

Qualifications

  • 10+ years of experience in large-scale distributed systems.
  • Hands-on experience with distributed training frameworks.
  • Strong understanding of training resilience patterns.
  • Solid grasp of GPU performance fundamentals.
  • Experience building multi-tenant platform products in the cloud.

Responsibilities

  • Drive the architecture and evolution of AIR.
  • Solve complex problems in large-scale training.
  • Push GPU efficiency and training performance.
  • Build resilience for multi-node jobs.
  • Collaborate with product, research, and platform teams.

Skills

Distributed systems
GPU training infrastructure
High-performance computing
ML systems
Distributed training frameworks
Communication skills
System design
Mentoring engineers

Education

BS in Computer Science or related field

Tools

PyTorch
FSDP
DeepSpeed
Megatron

Job description

Databricks is seeking a Staff Software Engineer to contribute to the AI Runtime platform for large-scale GPU training. The role centers around driving the architecture and scalability of the managed training stack, ensuring fast and reliable operations.

Candidates should have over 10 years of experience with distributed systems and GPU training, along with a passion for mentoring and engineering excellence. This position is based in San Francisco, California.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Runtime Architect - Scale GPU Training
Staff AI Runtime Architect - Scale GPU Training

Databricks • Mountain View (CA)

On-site
USD 190,000 - 265,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime Engineer: Scale GPU Training
Senior AI Runtime Engineer: Scale GPU Training

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,000 - 298,000
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,000 - 298,000
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (WY)

On-site
USD 229,000 - 297,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,000 - 298,000
Staff Software Engineer — Data & AI Platform (Equity)
Staff Software Engineer — Data & AI Platform (Equity)

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 192,000 - 260,000
Staff Software Engineer, AI Runtime
Staff Software Engineer, AI Runtime

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000