Staff GenAI Inference Engineer: Optimize LLM Serving Latency

Menlo Ventures

San Francisco (CA)

On-site

USD 190,900 - 232,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual performance bonus
Equity options
Comprehensive health benefits

Job summary

A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software engineering background and a proven ability to collaborate with researchers and drive architectural decisions. Competitive compensation is offered, with a salary range of $190,900 to $232,800 USD.

Qualifications

  • 6+ years of experience in performance-critical systems.
  • Proven track record of owning complex system components.
  • Hands-on experience with CUDA and GPU programming.

Responsibilities

  • Own and drive the architecture and implementation of the inference engine.
  • Lead the optimization for latency and throughput across GPUs.
  • Collaborate with researchers to integrate new model architectures.

Skills

CUDA programming
GPU programming
Distributed systems design
Communication skills
Performance optimization

Education

BS/MS/PhD in Computer Science

Tools

cuBLAS
cuDNN
NCCL

Job description

A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software engineering background and a proven ability to collaborate with researchers and drive architectural decisions. Competitive compensation is offered, with a salary range of $190,900 to $232,800 USD.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff GenAI Inference Architect: High-Throughput ML Serving
Staff GenAI Inference Architect: High-Throughput ML Serving

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Staff GenAI Kernel & Performance Engineer
Staff GenAI Kernel & Performance Engineer

Databricks • San Francisco (CA)

On-site
USD 190,000 - 233,000
Staff GenAI Inference Performance Engineer
Staff GenAI Inference Performance Engineer

Google • Mountain View (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Senior Inference Performance Engineer - GPU & CUDA
Senior Inference Performance Engineer - GPU & CUDA

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Senior Inference Performance Engineer — GPU & CUDA
Senior Inference Performance Engineer — GPU & CUDA

Inference • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
GenAI Inference Optimization Lead — GPU Performance
GenAI Inference Optimization Lead — GPU Performance

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 200,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Engineering Manager, GPU-Accelerated LLM Inference
Engineering Manager, GPU-Accelerated LLM Inference

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Distributed AI Inference Performance Engineer
Distributed AI Inference Performance Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer - Latency & Cost
Inference Performance Engineer - Latency & Cost

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity