Senior Software Engineer, Quantized Inference

NVIDIA Gruppe

Redmond (WA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe is seeking a Senior Software Engineer for Quantized Inference in Redmond, Washington. You will be responsible for implementing quantized and sparse recipes in inference engines and enhancing developer productivity across the team.

Successful candidates will be proficient in Python, possess strong software engineering fundamentals, and have experience with ML accelerators. This role will collaborate directly with partner inference teams to optimize productization efforts.

Qualifications

  • Proficient in Python and familiarity with C++ are required.
  • Strong software engineering fundamentals: concise, well-tested code.
  • Experience with ML accelerators and understanding of ML layer execution time.

Responsibilities

  • Implement quantized and sparse recipes in inference engines.
  • Own model export pipelines ensuring quantized checkpoints serialize correctly.
  • Build prototypes and benchmarking harnesses to evaluate recipe throughput.
  • Develop data analysis tooling and visualizations for numerics debugging.
  • Improve developer productivity across team processes.
  • Participate in code reviews and incorporate feedback.

Skills

Proficient in Python
Familiarity with C++
Strong software engineering fundamentals
Experience with ML accelerators

Job description

We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low‑precision or sparsified variants — unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines.

Each new recipe demands corresponding kernel and model‑level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per‑expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron‑LM, ModelOpt, and vLLM.

What you’ll be doing
  • Implement quantized and sparse recipes in inference engines (vLLM, TRT‑LLM, SGLang)
  • Own model export pipelines (ModelOpt, Megatron‑LM <=> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving
  • Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization
  • Develop data analysis tooling and visualizations for numerics debugging
  • Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction
  • Participate in code reviews and incorporate feedback
What we need to see
  • Proficient in Python; familiarity with C++
  • Strong software engineering fundamentals: concise, well‑tested code; fluent with AI‑assisted tooling
  • Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Quantized Inference
Senior Software Engineer, Quantized Inference

NVIDIA AI • Redmond (WA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Quantized Inference Engineer - Accelerate LLMs
Senior Quantized Inference Engineer - Accelerate LLMs

NVIDIA AI • Redmond (WA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Quantized Inference Engineer
Senior Quantized Inference Engineer

NVIDIA Gruppe • Redmond (WA)

On-site
USD 120,000 - 160,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
AI Engineer - Model Performance
AI Engineer - Model Performance

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive compensation and benefits
Supportive environment for innovation and growth
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior Inference Engineer
Senior Inference Engineer

Loft Labs, Inc. dba vCluster Labs • United States

On-site
USD 140,000 - 190,000