Staff Engineer, Foundation Model Serving & GPU Inference

Databricks

California (MO)

On-site

USD 192,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Databricks is seeking a Staff Engineer to design and build high-throughput, low-latency inference systems for the Foundation Model Serving API. You will shape core infrastructure, influence architecture, and collaborate across platform, product, and research teams to deliver a world-class model serving runtime.

As a senior contributor with 10+ years of distributed systems experience, you will mentor engineers, optimize GPU serving workloads, and help define long-term technical roadmap while

Qualifications

  • 10+ years of experience building and operating large-scale distributed systems.
  • Experience leading high-scale operationally sensitive backend systems.
  • Strong communication skills across fast-moving teams.
  • Strategic, product-oriented mindset with ability to align execution with long-term vision.

Responsibilities

  • Design and implement core systems and APIs powering Databricks Foundation Model Serving, ensuring scalability and reliability.
  • Define the technical roadmap and long-term architecture for serving workloads.
  • Drive architectural decisions to optimize performance, throughput, autoscaling for GPU serving workloads.
  • Mentor engineers through design reviews and technical guidance and collaborate across product, platform, and research teams.

Job description

Databricks is seeking a Staff Engineer to design and build high-throughput, low-latency inference systems for the Foundation Model Serving API. You will shape core infrastructure, influence architecture, and collaborate across platform, product, and research teams to deliver a world-class model serving runtime.

As a senior contributor with 10+ years of distributed systems experience, you will mentor engineers, optimize GPU serving workloads, and help define long-term technical roadmap while

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer: Scalable Model Serving & Inference
Staff Engineer: Scalable Model Serving & Inference

Databricks • California (MO)

On-site
USD 192,000 - 260,000
Senior Engineer, Model Serving & Low-Latency AI
Senior Engineer, Model Serving & Low-Latency AI

Databricks • California (MO)

On-site
USD 166,000 - 225,000
Staff Software Engineer: GenAI Inference & Scale
Staff Software Engineer: GenAI Inference & Scale

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer: Foundation Model API & GPU Inference
Staff Engineer: Foundation Model API & GPU Inference

Databricks Inc. • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Diversity and inclusion initiatives
Staff Engineer, Foundation Models & AI Infrastructure
Staff Engineer, Foundation Models & AI Infrastructure

Databricks • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Engineer - Foundation Model Serving & Inference
Staff Engineer - Foundation Model Serving & Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 120,000 - 160,000
Engineering Manager: Foundation Model Inference Leader
Engineering Manager: Foundation Model Inference Leader

Databricks • San Francisco (CA)

On-site
USD 190,000 - 262,000
Staff Software Engineer: Foundation Model Inference at Scale
Staff Software Engineer: Foundation Model Inference at Scale

Databricks • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Databricks Inc. • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Diversity and inclusion initiatives