Senior GenAI Kernel Performance Engineer

Databricks

California (MO)

On-site

USD 191,000 - 233,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Databricks is seeking a staff software engineer focusing on GenAI performance and kernel development. You will own high-performance GPU kernels for the GenAI inference stack, optimize for various hardware backends, and mentor teammates in kernel-level performance engineering.

You will collaborate with ML researchers, systems engineers, and product teams to push inference performance at scale, including profiling, optimization, and integration into production systems.

Qualifications

  • Hands-on kernel tuning for ML workloads (CUDA, Triton, OpenCL, LLVM IR).
  • Deep knowledge of GPU architectures and memory hierarchy.
  • Experience with advanced optimization: tiling, fusion, auto-tuning.
  • Familiarity with ML kernel libraries (cuBLAS, cuDNN, CUTLASS, oneDNN).
  • Strong debugging and profiling skills (Nsight, NVProf, perf, vtune).
  • Experience integrating optimized kernels into real-world ML inference systems.
  • Track record of shipping high-performance production software.

Responsibilities

  • Design, implement, benchmark, and maintain core compute kernels for GPU backends.
  • Drive kernel-level performance improvements: vectorization, tiling, fusion, auto-tuning.
  • Integrate kernel optimizations with higher-level ML systems.
  • Build profiling and verification tooling for correctness and performance.
  • Lead investigations of inference bottlenecks (memory bandwidth, cache).
  • Define reusable kernel abstractions and cross-backend portability.
  • Influence architecture decisions for memory layout and scheduling.
  • Mentor engineers and perform code reviews.
  • Collaborate with infra and ML teams to roll out optimizations.

Skills

Kernel tuning
CUDA
Triton
OpenCL
LLVM IR
GPU architecture
Profiling
ML inference
Distributed systems
Leadership

Education

BS/MS/PhD in Computer Science

Tools

Nsight
NVProf
perf
vtune
cuBLAS
cuDNN
CUTLASS
oneDNN

Job description

Databricks is seeking a staff software engineer focusing on GenAI performance and kernel development. You will own high-performance GPU kernels for the GenAI inference stack, optimize for various hardware backends, and mentor teammates in kernel-level performance engineering.

You will collaborate with ML researchers, systems engineers, and product teams to push inference performance at scale, including profiling, optimization, and integration into production systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Senior GenAI Kernel & GPU Optimization Engineer
Senior GenAI Kernel & GPU Optimization Engineer

Databricks • Mountain View (CA)

On-site
USD 166,000 - 225,000
Comprehensive benefits
Annual performance bonus
Equity options
Staff Software Engineer: GenAI Inference & Scale
Staff Software Engineer: GenAI Inference & Scale

Databricks • California (MO)

On-site
USD 191,000 - 233,000
Senior GPU Kernel Performance Engineer
Senior GPU Kernel Performance Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 140,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+2
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Senior GPU Software Engineer — AI & Kernel Performance
Senior GPU Software Engineer — AI & Kernel Performance

AMD • United States

On-site
USD 150,000 - 210,000
Senior CUDA Kernel & Performance Engineer
Senior CUDA Kernel & Performance Engineer

General Motors • Washington

Hybrid
USD 170,000 - 258,000
Health and wellbeing benefits
Hybrid work option
Competitive compensation package
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior AI Kernel Engineer — Remote/Hybrid Inference
Senior AI Kernel Engineer — Remote/Hybrid Inference

Modular • United States

Hybrid
USD 198,000 - 286,000
Amazing Team
World-class Benefits
Competitive Compensation
+1