GenAI Inference Software Engineer — Scalable LLM Serving

Menlo Ventures

San Francisco (CA)

On-site

USD 142,200 - 204,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Databricks is seeking a Software Engineer for GenAI inference in San Francisco. You will design and optimize the inference engine powering our Foundation Model API. Responsibilities include collaborating across teams, enhancing performance for large-scale ML models, and implementing new features. Ideal candidates should possess a strong software engineering background, understanding of ML inference, and experience with CUDA. This role offers a salary range of $142,200 to $204,600, alongside benefits and opportunities for advancement.

Qualifications

  • Strong software engineering background, 3+ years in performance-critical systems.
  • Hands-on experience with CUDA and GPU programming.
  • Ability to work closely with ML researchers to bring novel ideas into production.

Responsibilities

  • Design and implement the inference engine optimized for large-scale LLMs.
  • Collaborate with researchers on new model architectures.
  • Optimize for latency, throughput, and memory efficiency across GPUs.

Skills

Performance-critical systems
ML inference internals
CUDA
GPU programming
Distributed systems design
Performance bottlenecks
Instrumentation and profiling tools

Education

BS/MS/PhD in Computer Science or related field

Tools

cuBLAS
cuDNN
NCCL

Job description

Databricks is seeking a Software Engineer for GenAI inference in San Francisco. You will design and optimize the inference engine powering our Foundation Model API. Responsibilities include collaborating across teams, enhancing performance for large-scale ML models, and implementing new features. Ideal candidates should possess a strong software engineering background, understanding of ML inference, and experience with CUDA. This role offers a salary range of $142,200 to $204,600, alongside benefits and opportunities for advancement.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GenAI Inference Architect: High-Throughput ML Serving
Staff GenAI Inference Architect: High-Throughput ML Serving

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
GenAI Inference Engineer — High-Performance ML Systems
GenAI Inference Engineer — High-Performance ML Systems

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 142,000 - 205,000
Staff GenAI Inference Engineer: Optimize LLM Serving Latency
Staff GenAI Inference Engineer: Optimize LLM Serving Latency

Menlo Ventures • San Francisco (CA)

On-site
USD 190,000 - 233,000
Annual performance bonus
Equity options
Comprehensive health benefits
Staff Software Engineer, Inference Systems at Scale
Staff Software Engineer, Inference Systems at Scale

Anthropic • San Francisco (CA)

On-site
USD 300,000 - 485,000
Competitive salary
Flexible working hours
Generous vacation and parental leave
Staff Software Engineer - GenAI inference
Staff Software Engineer - GenAI inference

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Comprehensive benefits and perks
Annual performance bonus
Equity options
Staff GenAI ML Engineer — Build Next-Gen AI at Scale
Staff GenAI ML Engineer — Build Next-Gen AI at Scale

Databricks • San Francisco (CA)

On-site
USD 190,000 - 285,000
Distributed LLM Inference Engineer - Scale AI at Speed
Distributed LLM Inference Engineer - Scale AI at Speed

Anyscale • San Francisco (CA)

On-site
USD 120,000 - 160,000
Stock Options
Healthcare plans
401k Retirement Plan
+6
Engineering Manager, ML Inference & Scale
Engineering Manager, ML Inference & Scale

Anthropic • San Francisco (CA)

Hybrid
USD 425,000 - 560,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Software Engineer - GenAI inference
Software Engineer - GenAI inference

Cacheflow • San Francisco (CA)

On-site
USD 142,000 - 205,000
AI Inference Engineer: Real-Time ML, Hybrid, Equity
AI Inference Engineer: Real-Time ML, Hybrid, Equity

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1