Senior ML Engineer: Quantized Inference & Pipelines

NVIDIA AI

Redmond (WA)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA AI is seeking an engineer to implement quantized and sparse recipes in inference engines and to manage model export pipelines for correct serialization. You will build benchmarking harnesses and data analysis tools to improve developer productivity through infrastructure and CI improvements.

The role requires strong Python and C++ skills, experience with ML accelerators, and familiarity with PyTorch internals, with 4+ years in software engineering. MS/PhD in CS is preferred.

Qualifications

  • Requires 4+ years in a relevant software engineering role.
  • Proficiency in Python and familiarity with C++, plus strong fundamentals.

Responsibilities

  • Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization.
  • Develop benchmarking harnesses and data analysis tools to improve performance and visibility.
  • Improve developer productivity through infrastructure and CI improvements.

Skills

Python
C++
Triton Kernels
PyTorch
Quantized Inference
Model Compression
Machine Learning Accelerators
vLLM
TRT-LLM
SGLang
Megatron-LM
ModelOpt
Software Engineering
Data Analysis
Numerical Debugging
Large Language Models

Education

Master's degree in Computer Science
PhD in Computer Science

Tools

PyTorch internals

Job description

NVIDIA AI is seeking an engineer to implement quantized and sparse recipes in inference engines and to manage model export pipelines for correct serialization. You will build benchmarking harnesses and data analysis tools to improve developer productivity through infrastructure and CI improvements.

The role requires strong Python and C++ skills, experience with ML accelerators, and familiarity with PyTorch internals, with 4+ years in software engineering. MS/PhD in CS is preferred.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Quantized Inference
Senior Software Engineer, Quantized Inference

NVIDIA AI • Redmond (WA)

On-site
USD 140,000 - 200,000
Equity
Benefits
Senior ML Quantization Engineer for Optical AI Compute
Senior ML Quantization Engineer for Optical AI Compute

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 240,000
Health coverage
HSA contributions
Unlimited PTO
+3
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Systems Engineer: GPU Kernels & Inference Equity
Senior AI Systems Engineer: GPU Kernels & Inference Equity

NVIDIA AI • Michigan

On-site
USD 150,000 - 190,000
Equity
Health Insurance
Senior Applied Scientist — Optical AI Quantization
Senior Applied Scientist — Optical AI Quantization

Neurophos • Austin (TX)

On-site
USD 180,000 - 260,000
Health plan premiums coverage
Unlimited PTO
401(k) matching
+2
Machine Learning Engineer
Machine Learning Engineer

Neutral Street • New York (NY)

On-site
USD 120,000 - 180,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA • Westford (MA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000