Senior Engineer, Model Serving & Low-Latency AI

Databricks

California (MO)

On-site

USD 166,000 - 225,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Databricks is seeking a Senior Engineer to shape the Model Serving platform and its infrastructure, enabling high-throughput, low-latency inference across CPU and GPU workloads.

You will design core systems, guide architectural decisions, and collaborate with platform, product, infrastructure, and research teams to deliver scalable, reliable serving capabilities while mentoring teammates and promoting engineering excellence.

Qualifications

  • 5+ years of experience building and operating large-scale distributed systems.
  • Experience in model serving, inference systems, or related infrastructure (e.g., routing, scheduling, autoscaling, observability).
  • Strong foundation in algorithms, data structures, and system design as applied to large-scale, low-latency serving systems.

Responsibilities

  • Design and implement core systems and APIs powering Databricks Model Serving, ensuring scalability and reliability.
  • Drive architectural decisions to optimize performance, throughput, and operational efficiency for CPU/GPU workloads.
  • Contribute to components across the serving stack—from model container builds to routing, caching, and observability.

Skills

Distributed systems
Model serving / inference
Algorithms & data structures
Mentoring / leadership
Cross-functional collaboration

Job description

Databricks is seeking a Senior Engineer to shape the Model Serving platform and its infrastructure, enabling high-throughput, low-latency inference across CPU and GPU workloads.

You will design core systems, guide architectural decisions, and collaborate with platform, product, infrastructure, and research teams to deliver scalable, reliable serving capabilities while mentoring teammates and promoting engineering excellence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer: Scalable Model Serving & Inference
Staff Engineer: Scalable Model Serving & Inference

Databricks • California (MO)

On-site
USD 192,000 - 260,000
Staff Engineer, Foundation Model Serving & GPU Inference
Staff Engineer, Foundation Model Serving & GPU Inference

Databricks • California (MO)

On-site
USD 192,000 - 260,000
Senior AI Infra Engineer — Inference Platform
Senior AI Infra Engineer — Inference Platform

Databricks Inc. • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Software Engineer: GenAI Inference & Scale
Staff Software Engineer: GenAI Inference & Scale

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior Engineer, Model Serving & Inference
Senior Engineer, Model Serving & Inference

Databricks • San Francisco (CA)

On-site
USD 166,000 - 225,000
Senior Software Engineer, Model Serving
Senior Software Engineer, Model Serving

Databricks • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior Backend Engineer, AI Platform & Infra
Senior Backend Engineer, AI Platform & Infra

Databricks • California (MO)

On-site
USD 166,000 - 225,000
Senior Software Engineer, Model Serving
Senior Software Engineer, Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package