Staff GenAI Inference Architect: High-Throughput ML Serving
Databricks
San Francisco (CA)
On-site
USD 190,900 - 232,800
Full time
14 days+
Application generator
A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Benefits offered by this job
Comprehensive benefits and perks
Annual performance bonus
Equity options
Job summary
Databricks is seeking a Staff Software Engineer for GenAI inference to drive the architecture, development, and optimization of their inference engine. The role involves high-throughput, low-latency solutions and full GenAI inference stack management. Candidates must have a strong software engineering background (6+ years), knowledge in ML inference, as well as CUDA and distributed systems design skills. Competitive pay range is $190,900—$232,800, including potential bonuses and benefits.
Qualifications
6+ years of experience in performance-critical systems.
Proven track record of driving architectural decisions.
Hands-on experience with CUDA and GPU programming.
Responsibilities
Lead the architecture and optimization of the inference engine.
Collaborate with researchers to implement new model features.
Ensure reliability and fault tolerance in inference pipelines.
Skills
Software engineering background
ML inference internals
CUDA and GPU programming
Distributed systems design
Performance optimization
Education
BS/MS/PhD in Computer Science or related field
Tools
cuBLAS
cuDNN
NCCL
Job description
Databricks is seeking a Staff Software Engineer for GenAI inference to drive the architecture, development, and optimization of their inference engine. The role involves high-throughput, low-latency solutions and full GenAI inference stack management. Candidates must have a strong software engineering background (6+ years), knowledge in ML inference, as well as CUDA and distributed systems design skills. Competitive pay range is $190,900—$232,800, including potential bonuses and benefits.