Staff GenAI Inference Architect: High-Throughput ML Serving

Databricks

San Francisco (CA)

On-site

USD 190,900 - 232,800

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive benefits and perks
Annual performance bonus
Equity options

Job summary

Databricks is seeking a Staff Software Engineer for GenAI inference to drive the architecture, development, and optimization of their inference engine. The role involves high-throughput, low-latency solutions and full GenAI inference stack management. Candidates must have a strong software engineering background (6+ years), knowledge in ML inference, as well as CUDA and distributed systems design skills. Competitive pay range is $190,900—$232,800, including potential bonuses and benefits.

Qualifications

  • 6+ years of experience in performance-critical systems.
  • Proven track record of driving architectural decisions.
  • Hands-on experience with CUDA and GPU programming.

Responsibilities

  • Lead the architecture and optimization of the inference engine.
  • Collaborate with researchers to implement new model features.
  • Ensure reliability and fault tolerance in inference pipelines.

Skills

Software engineering background
ML inference internals
CUDA and GPU programming
Distributed systems design
Performance optimization

Education

BS/MS/PhD in Computer Science or related field

Tools

cuBLAS
cuDNN
NCCL

Job description

Databricks is seeking a Staff Software Engineer for GenAI inference to drive the architecture, development, and optimization of their inference engine. The role involves high-throughput, low-latency solutions and full GenAI inference stack management. Candidates must have a strong software engineering background (6+ years), knowledge in ML inference, as well as CUDA and distributed systems design skills. Competitive pay range is $190,900—$232,800, including potential bonuses and benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GenAI Inference Engineer: Optimize LLM Serving Latency
Staff GenAI Inference Engineer: Optimize LLM Serving Latency

Menlo Ventures • San Francisco (CA)

On-site
USD 190,000 - 233,000
Annual performance bonus
Equity options
Comprehensive health benefits
Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Comprehensive benefits and perks
Annual performance bonus
Equity options
Staff Software Engineer: GenAI Inference & Scale
Staff Software Engineer: GenAI Inference & Scale

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff GenAI Inference Performance Engineer
Staff GenAI Inference Performance Engineer

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Staff AI Inference Performance Engineer (GenAI)
Staff AI Inference Performance Engineer (GenAI)

Google • Town of Montana (WI)

On-site
USD 207,000 - 300,000
Bonus target
Equity
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Senior GenAI ML Engineer - High-Performance Inference
Senior GenAI ML Engineer - High-Performance Inference

Adobe • San Jose (CA)

On-site
USD 183,000 - 266,000
Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000