Senior AI Runtime Engineer: Scale GPU Training

Cacheflow

San Francisco (CA)

On-site

USD 160,000 - 225,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Annual performance bonus
Comprehensive benefits package

Job summary

Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of accelerators.

The role demands 5+ years of experience in distributed systems, proficiency with GPU training infrastructure, and a commitment to driving engineering excellence. The pay range is between $160,000 and $225,000 USD.

Qualifications

  • 5+ years of experience building and operating large-scale distributed systems.
  • Strong understanding of training resilience patterns.
  • Proven ability to deliver technically complex projects.

Responsibilities

  • Drive the architecture and evolution of AIR's managed GPU training platform.
  • Solve large-scale training challenges, including orchestration and performance.
  • Champion engineering excellence and mentor engineers.

Skills

Large-scale distributed systems
GPU training infrastructure
High-performance computing
ML systems
Distributed training frameworks
Algorithms and data structures
System design

Education

BS in Computer Science or related field
MS or PhD preferred

Job description

Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of accelerators.

The role demands 5+ years of experience in distributed systems, proficiency with GPU training infrastructure, and a commitment to driving engineering excellence. The pay range is between $160,000 and $225,000 USD.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Runtime Engineer — GPU Training Platform
Senior AI Runtime Engineer — GPU Training Platform

Databricks • Mountain View (CA)

On-site
USD 160,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Senior AI Runtime & GPU Training Systems Engineer
Senior AI Runtime & GPU Training Systems Engineer

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 265,000
Equity
Annual performance bonus
Comprehensive benefits
Staff Software Engineer, AI Runtime - Scalable GPU Training
Staff Software Engineer, AI Runtime - Scalable GPU Training

Databricks Inc. • Mountain View (CA)

On-site
USD 190,000 - 265,000
Senior AI Runtime Engineering Manager: Scale GPU Training
Senior AI Runtime Engineering Manager: Scale GPU Training

Menlo Ventures • San Francisco (CA)

On-site
USD 228,600 - 297,120
Senior Engineering Manager, AI Runtime & GPU Training
Senior Engineering Manager, AI Runtime & GPU Training

Databricks • Mountain View (CA)

On-site
USD 228,600 - 297,120
Staff AI Infra Engineer — Scale GPU Experiments
Staff AI Infra Engineer — Scale GPU Experiments

Databricks • San Francisco (CA)

On-site
USD 190,000 - 270,000
Senior Software Engineer, AI Runtime
Senior Software Engineer, AI Runtime

Cacheflow • San Francisco (CA)

On-site
USD 160,000 - 225,000
Equity
Annual performance bonus
Comprehensive benefits package
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Staff AI Research Infra Engineer - GPU Scale & Equity
Staff AI Research Infra Engineer - GPU Scale & Equity

Databricks Inc. • New York (NY)

On-site
USD 199,000 - 270,000
Senior AI Runtime Platform Leader
Senior AI Runtime Platform Leader

jobr.pro • Mountain View (CA)

On-site
USD 228,600 - 297,120