Staff Software Engineer: GenAI Inference & Scale

Databricks

California (MO)

On-site

USD 191,000 - 233,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Databricks is seeking a Staff Software Engineer to lead the GenAI inference stack for their Foundation Model API. You will own the architecture, design, and optimization of the inference engine, targeting high throughput, low latency, and scalable performance across GPUs and accelerators.

You will collaborate with researchers to deploy new model architectures, drive optimizations across kernels, runtimes, and orchestration, and ensure reliability with instrumentation and monitoring.

Qualifications

  • BS/MS/PhD in CS or related field.
  • 6+ years in performance-critical systems.
  • Experience owning complex components and driving architecture end-to-end.
  • Deep ML inference internals: attention, MLPs, etc.
  • Hands-on CUDA/GPU programming with libraries (cuBLAS/cuDNN/NCCL).
  • Distributed systems design including RPC, queues, batching, sharding, memory partitioning.
  • Experience building instrumentation, tracing, profiling tools for ML.
  • Ability to lead through influence with researchers and production teams.

Responsibilities

  • Own architecture, design, and implementation of the inference engine for large-scale LLMs.
  • Partner with researchers to bring new model architectures into the engine.
  • Lead end-to-end optimization for latency, throughput, memory, and hardware use.
  • Define standards for instrumentation, profiling, and tracing tooling.
  • Architect scalable routing, batching, scheduling, and memory management for inference workloads.
  • Ensure reliability, reproducibility, and fault tolerance in inference pipelines.

Skills

Performance-critical systems
CUDA programming
Distributed systems design
Instrumentation & profiling
Leadership & collaboration

Education

BS/MS/PhD in Computer Science or related field

Tools

CUDA
cuBLAS
cuDNN
NCCL

Job description

Databricks is seeking a Staff Software Engineer to lead the GenAI inference stack for their Foundation Model API. You will own the architecture, design, and optimization of the inference engine, targeting high throughput, low latency, and scalable performance across GPUs and accelerators.

You will collaborate with researchers to deploy new model architectures, drive optimizations across kernels, runtimes, and orchestration, and ensure reliability with instrumentation and monitoring.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, Foundation Model Serving & GPU Inference
Staff Engineer, Foundation Model Serving & GPU Inference

Databricks • California (MO)

On-site
USD 192,000 - 260,000
Staff Engineer: Scalable Model Serving & Inference
Staff Engineer: Scalable Model Serving & Inference

Databricks • California (MO)

On-site
USD 192,000 - 260,000
Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Comprehensive benefits and perks
Annual performance bonus
Equity options
Staff GenAI Inference Architect: High-Throughput ML Serving
Staff GenAI Inference Architect: High-Throughput ML Serving

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Senior GenAI Kernel Performance Engineer
Senior GenAI Kernel Performance Engineer

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Cacheflow • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Infra Engineer — Inference Platform
Senior AI Infra Engineer — Inference Platform

Databricks Inc. • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff GenAI Engineer - Scale AI Features
Staff GenAI Engineer - Scale AI Features

ACM CAIS 2026 • San Francisco (CA)

On-site
USD 190,000 - 285,000
Staff Software Engineer: Foundation Model Inference at Scale
Staff Software Engineer: Foundation Model Inference at Scale

Databricks • San Francisco (CA)

On-site
USD 190,000 - 265,000