Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential

United States

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Confidential in the United States seeks a senior engineer to own performance optimization for production LLM inference, focusing on latency, throughput, and cost across kernels and serving engines. You will profile GPU performance, apply quantization and batching strategies, and extend serving stacks like vLLM, TensorRT-LLM, and Triton.

You will collaborate with model and platform teams to push new architectures from works to fast, while targeting multi-GPU and accelerator-rich deployments.

Qualifications

  • Deep experience optimizing deep-learning inference in production.
  • Hands-on GPU programming and performance engineering (CUDA or equivalent).
  • Fluency with modern LLM serving stacks (vLLM / TensorRT-LLM / SGLang / Triton).
  • A track record of measurable performance wins (latency / throughput / cost).
  • Strong systems fundamentals and a profiling-first mindset.

Responsibilities

  • Optimize LLM inference for latency, throughput, and cost — at the kernel and serving-engine level.
  • Profile and tune GPU performance (CUDA, TensorRT-LLM); apply quantization, speculative decoding, and batching strategies.
  • Get the most out of serving frameworks like vLLM, SGLang, and Triton — and extend them where they fall short.
  • Optimize across hardware targets where relevant (NVIDIA and other accelerators).
  • Partner with model and platform teams to take new architectures from "works" to "fast".

Skills

Deep learning inference
GPU programming
Performance engineering
Profiling
CUDA

Tools

TensorRT-LLM
vLLM
SGLang
Triton

Job description

Own the performance of large language models in production — the latency, the throughput, the cost-per-token. This is deep inference-optimization work: profiling and tuning at the GPU and serving-engine level to make models run faster and cheaper at scale. You'll join a small, senior team at an established enterprise software company building LLM-powered capabilities into its products.

What you'll do:

  • Optimize LLM inference for latency, throughput, and cost — at the kernel and serving-engine level
  • Profile and tune GPU performance (CUDA, TensorRT-LLM); apply quantization, speculative decoding, and batching strategies
  • Get the most out of serving frameworks like vLLM, SGLang, and Triton — and extend them where they fall short
  • Optimize across hardware targets where relevant (NVIDIA and other accelerators)
  • Partner with model and platform teams to take new architectures from "works" to "fast"

What you'll bring:

  • Deep experience optimizing deep-learning inference in production
  • Hands-on GPU programming and performance engineering (CUDA or equivalent)
  • Fluency with modern LLM serving stacks (vLLM / TensorRT-LLM / SGLang / Triton)
  • A track record of measurable performance wins (latency / throughput / cost)
  • Strong systems fundamentals and a profiling-first mindset

Nice to have:

  • Kernel-level contributions to open-source inference projects
  • Experience across multiple accelerator types
  • Distributed / multi-GPU serving experience

A rare role where deep performance work is the whole job, not a side quest.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Infrastructure Engineer, LLM Inference Optimization
Infrastructure Engineer, LLM Inference Optimization

GMI Cloud • Mountain View (CA)

On-site
USD 170,000 - 230,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits