Senior Quantized Inference Engineer - Accelerate LLMs

NVIDIA AI

Redmond (WA)

On-site

USD 184,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams to optimize throughput and interactivity across Megatron-LM, ModelOpt, and vLLM.

The role requires strong Python skills with familiarity in C++, experience with PyTorch internals, and 4+ years in software engineering.

Qualifications

  • 4+ years in a relevant software engineering role.
  • MS/PhD in Computer Science or related field or equivalent experience.
  • Familiarity with PyTorch internals (custom ops, autograd, export) or equivalent framework.
  • Experience with ML accelerators and quantization concepts is a plus.

Responsibilities

  • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang).
  • Own model export pipelines (ModelOpt, Megatron-LM HuggingFace).
  • Build prototypes and benchmarking harnesses for throughput and interactivity.
  • Develop data analysis tooling and visualizations for numerics debugging.
  • Improve developer productivity: CI, build systems, training infrastructure.
  • Participate in code reviews and incorporate feedback.

Skills

Python
C++
PyTorch internals
ML accelerators

Education

MS/PhD in Computer Science or related field

Job description

NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams to optimize throughput and interactivity across Megatron-LM, ModelOpt, and vLLM.

The role requires strong Python skills with familiarity in C++, experience with PyTorch internals, and 4+ years in software engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Quantized Inference
Senior Software Engineer, Quantized Inference

NVIDIA Gruppe • Redmond (WA)

On-site
USD 120,000 - 160,000
Senior Software Engineer, Quantized Inference
Senior Software Engineer, Quantized Inference

NVIDIA AI • Redmond (WA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Quantized Inference Engineer
Senior Quantized Inference Engineer

NVIDIA Gruppe • Redmond (WA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Algorithms Engineer — Equity Options
Senior LLM Inference Algorithms Engineer — Equity Options

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior Applied Scientist — Efficient LLM Inference & Optimization
Senior Applied Scientist — Efficient LLM Inference & Optimization

Nebius • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Senior ML Inference Engineer – TensorRT & LLMs
Senior ML Inference Engineer – TensorRT & LLMs

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000