Senior LLM Inference Engineer: Performance & Optimization

Confidential

United States

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Confidential in the United States seeks a senior engineer to own performance optimization for production LLM inference, focusing on latency, throughput, and cost across kernels and serving engines. You will profile GPU performance, apply quantization and batching strategies, and extend serving stacks like vLLM, TensorRT-LLM, and Triton.

You will collaborate with model and platform teams to push new architectures from works to fast, while targeting multi-GPU and accelerator-rich deployments.

Qualifications

  • Deep experience optimizing deep-learning inference in production.
  • Hands-on GPU programming and performance engineering (CUDA or equivalent).
  • Fluency with modern LLM serving stacks (vLLM / TensorRT-LLM / SGLang / Triton).
  • A track record of measurable performance wins (latency / throughput / cost).
  • Strong systems fundamentals and a profiling-first mindset.

Responsibilities

  • Optimize LLM inference for latency, throughput, and cost — at the kernel and serving-engine level.
  • Profile and tune GPU performance (CUDA, TensorRT-LLM); apply quantization, speculative decoding, and batching strategies.
  • Get the most out of serving frameworks like vLLM, SGLang, and Triton — and extend them where they fall short.
  • Optimize across hardware targets where relevant (NVIDIA and other accelerators).
  • Partner with model and platform teams to take new architectures from "works" to "fast".

Skills

Deep learning inference
GPU programming
Performance engineering
Profiling
CUDA

Tools

TensorRT-LLM
vLLM
SGLang
Triton

Job description

Confidential in the United States seeks a senior engineer to own performance optimization for production LLM inference, focusing on latency, throughput, and cost across kernels and serving engines. You will profile GPU performance, apply quantization and batching strategies, and extend serving stacks like vLLM, TensorRT-LLM, and Triton.

You will collaborate with model and platform teams to push new architectures from works to fast, while targeting multi-GPU and accelerator-rich deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Remote LLM Inference Optimization Lead
Senior Remote LLM Inference Optimization Lead

Dragonfly Digital Management, LLC (Dragonfly Capital) • United States

On-site
USD 140,000 - 210,000
Senior LLM Inference Architect — Performance & Benchmarking
Senior LLM Inference Architect — Performance & Benchmarking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Training Performance Architect
Senior LLM Training Performance Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 356,500
Equity
Benefits