Databricks is seeking a Software Engineer for GenAI inference in San Francisco. You will design and optimize the inference engine powering our Foundation Model API. Responsibilities include collaborating across teams, enhancing performance for large-scale ML models, and implementing new features. Ideal candidates should possess a strong software engineering background, understanding of ML inference, and experience with CUDA. This role offers a salary range of $142,200 to $204,600, alongside benefits and opportunities for advancement.
Qualifications
Strong software engineering background, 3+ years in performance-critical systems.
Hands-on experience with CUDA and GPU programming.
Ability to work closely with ML researchers to bring novel ideas into production.
Responsibilities
Design and implement the inference engine optimized for large-scale LLMs.
Collaborate with researchers on new model architectures.
Optimize for latency, throughput, and memory efficiency across GPUs.
Skills
Performance-critical systems
ML inference internals
CUDA
GPU programming
Distributed systems design
Performance bottlenecks
Instrumentation and profiling tools
Education
BS/MS/PhD in Computer Science or related field
Tools
cuBLAS
cuDNN
NCCL
Job description
Databricks is seeking a Software Engineer for GenAI inference in San Francisco. You will design and optimize the inference engine powering our Foundation Model API. Responsibilities include collaborating across teams, enhancing performance for large-scale ML models, and implementing new features. Ideal candidates should possess a strong software engineering background, understanding of ML inference, and experience with CUDA. This role offers a salary range of $142,200 to $204,600, alongside benefits and opportunities for advancement.