Staff Engineer: Scalable Model Serving & Inference

Databricks

California (MO)

On-site

USD 192,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Databricks is seeking a Staff Engineer to shape the Model Serving platform, designing scalable systems and APIs for high-throughput, low-latency inference on CPU and GPU. You will influence architectural direction and collaborate across platform, product, infrastructure, and research teams to deliver a world-class serving platform.

You will lead architectural decisions, drive performance optimizations, and mentor engineers, while ensuring operational excellence and adherence to best practices in

Qualifications

  • 10+ years of experience building and operating large-scale distributed systems.
  • Deep expertise in model serving, inference systems, and related infrastructure (routing, scheduling, autoscaling, observability).
  • Strong foundation in algorithms, data structures, and system design for low-latency serving systems.

Responsibilities

  • Design and implement core systems and APIs powering Databricks Model Serving, ensuring scalability and reliability.
  • Define the technical roadmap and long-term architecture for serving workloads.
  • Drive architectural decisions to optimize performance, throughput, and autoscaling for CPU and GPU workloads.
  • Contribute to components across the serving infrastructure (model containers, deployment workflows, routing, caching, observability, autoscaling).
  • Collaborate with product, platform, and research teams to translate customer needs into reliable systems.
  • Lead initiatives to improve latency, availability, and cost-effectiveness across serving layers.
  • Establish best practices for code quality, testing, and operational readiness; mentor engineers.

Skills

Distributed systems
Model serving
Inference systems
System design
CPU/GPU inference
Architectural direction
Leadership
Mentoring

Job description

Databricks is seeking a Staff Engineer to shape the Model Serving platform, designing scalable systems and APIs for high-throughput, low-latency inference on CPU and GPU. You will influence architectural direction and collaborate across platform, product, infrastructure, and research teams to deliver a world-class serving platform.

You will lead architectural decisions, drive performance optimizations, and mentor engineers, while ensuring operational excellence and adherence to best practices in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineer, Model Serving & Low-Latency AI
Senior Engineer, Model Serving & Low-Latency AI

Databricks • California (MO)

On-site
USD 166,000 - 225,000
Staff Engineer, Foundation Model Serving & GPU Inference
Staff Engineer, Foundation Model Serving & GPU Inference

Databricks • California (MO)

On-site
USD 192,000 - 260,000
Staff Software Engineer: GenAI Inference & Scale
Staff Software Engineer: GenAI Inference & Scale

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Infra Engineer — Inference Platform
Senior AI Infra Engineer — Inference Platform

Databricks Inc. • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Remote Model Serving Engineer - Scale ML Inference
Remote Model Serving Engineer - Scale ML Inference

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 74,000 - 98,000
Staff Software Engineer, Model Serving
Staff Software Engineer, Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Eligibility for annual performance bonus
Equity opportunities
Staff Software Engineer: Foundation Model Inference at Scale
Staff Software Engineer: Foundation Model Inference at Scale

Databricks • San Francisco (CA)

On-site
USD 190,000 - 265,000
Senior Software Engineer, Model Serving
Senior Software Engineer, Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package