Staff Software Engineer - GenAI inference

Databricks

San Francisco (CA)

On-site

USD 190,900 - 232,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive benefits and perks
Annual performance bonus
Equity options

Job summary

Databricks is seeking a Staff Software Engineer for GenAI inference to drive the architecture, development, and optimization of their inference engine. The role involves high-throughput, low-latency solutions and full GenAI inference stack management. Candidates must have a strong software engineering background (6+ years), knowledge in ML inference, as well as CUDA and distributed systems design skills. Competitive pay range is $190,900—$232,800, including potential bonuses and benefits.

Qualifications

  • 6+ years of experience in performance-critical systems.
  • Proven track record of driving architectural decisions.
  • Hands-on experience with CUDA and GPU programming.

Responsibilities

  • Lead the architecture and optimization of the inference engine.
  • Collaborate with researchers to implement new model features.
  • Ensure reliability and fault tolerance in inference pipelines.

Skills

Software engineering background
ML inference internals
CUDA and GPU programming
Distributed systems design
Performance optimization

Education

BS/MS/PhD in Computer Science or related field

Tools

cuBLAS
cuDNN
NCCL

Job description

About This Role

As a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers the Databricks Foundation Model API. You’ll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems.

P-1285

What You Will Do
  • Own and drive the architecture, design, and implementation of the inference engine, and collaborate on model‑serving stack optimized for large‑scale LLMs inference
  • Partner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
  • Lead the end‑to‑end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
  • Define and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
  • Architect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads
  • Ensure reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning
  • Collaborate cross‑functionally on integrating with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
  • Drive cross‑team collaboration: with platform engineers, cloud infrastructure, and security/compliance teams
  • Represent the team externally through benchmarks, whitepapers, and open‑source contributions
What We Look For
  • BS/MS/PhD in Computer Science, or a related field
  • Strong software engineering background (6+ years or equivalent) in performance‑critical systems
  • Proven track record of owning complex system components and driving architectural decisions end‑to‑end
  • Deep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.
  • Hands‑on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
  • Strong background in distributed systems design, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
  • Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)
  • Experience building instrumentation, tracing, and profiling tools for ML models
  • Ability to lead through influence – work closely with ML researchers, translate novel model ideas into production systems
  • Excellent communication and leadership skills, with a proactive and ownership‑driven mindset
  • Bonus: published research or open‑source contributions in ML systems, inference optimization, or model serving
Pay Range Transparency

Local Pay Range: $190,900—$232,800 USD. Compensation may include eligibility for annual performance bonus, equity, and benefits.

Benefits

Comprehensive benefits and perks are offered.

Our Commitment to Diversity and Inclusion

We are committed to fostering a diverse and inclusive culture and have hiring practices that meet equal employment opportunity standards. Disparate qualities are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio‑economic status, veteran status, and other protected characteristics.

Compliance

If access to export‑controlled technology or source code is required for performance of job duties, the Employer may decline to proceed with an applicant on this basis alone.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 233,000
Annual performance bonus
Equity options
Comprehensive health benefits
Software Engineer - GenAI inference
Software Engineer - GenAI inference

Cacheflow • San Francisco (CA)

On-site
USD 142,000 - 205,000
Software Engineer - GenAI inference Databricks San Francisco, California
Software Engineer - GenAI inference Databricks San Francisco, California

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 142,000 - 205,000
Staff Software Engineer- Foundation Model Inference San Francisco, California
Staff Software Engineer- Foundation Model Inference San Francisco, California

Databricks Inc. • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Software Engineer - GenAI Performance and Kernel
Staff Software Engineer - GenAI Performance and Kernel

Cacheflow • San Francisco (CA)

On-site
USD 190,000 - 233,000
Comprehensive benefits
Annual performance bonus
Equity options
Staff Software Engineer, Foundational Model Serving
Staff Software Engineer, Foundational Model Serving

Databricks Inc. • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Diversity and inclusion initiatives
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Backend Software Engineer- (AI Platform)
Staff Backend Software Engineer- (AI Platform)

Menlo Ventures • San Francisco (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Annual performance bonus
Equity options
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

Databricks • San Francisco (CA)

On-site
USD 190,000 - 265,000
Staff Software Engineer, Foundational Model Serving Databricks San Francisco, California
Staff Software Engineer, Foundational Model Serving Databricks San Francisco, California

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 192,000 - 260,000