Senior Software Engineer, Quantized Inference

NVIDIA AI

Redmond (WA)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA AI is seeking an engineer to implement quantized and sparse recipes in inference engines and to manage model export pipelines for correct serialization. You will build benchmarking harnesses and data analysis tools to improve developer productivity through infrastructure and CI improvements.

The role requires strong Python and C++ skills, experience with ML accelerators, and familiarity with PyTorch internals, with 4+ years in software engineering. MS/PhD in CS is preferred.

Qualifications

  • Requires 4+ years in a relevant software engineering role.
  • Proficiency in Python and familiarity with C++, plus strong fundamentals.

Responsibilities

  • Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization.
  • Develop benchmarking harnesses and data analysis tools to improve performance and visibility.
  • Improve developer productivity through infrastructure and CI improvements.

Skills

Python
C++
Triton Kernels
PyTorch
Quantized Inference
Model Compression
Machine Learning Accelerators
vLLM
TRT-LLM
SGLang
Megatron-LM
ModelOpt
Software Engineering
Data Analysis
Numerical Debugging
Large Language Models

Education

Master's degree in Computer Science
PhD in Computer Science

Tools

PyTorch internals

Job description

Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization. Develop benchmarking harnesses, data analysis tools, and improve developer productivity through infrastructure and CI improvements.

Requirements:

Requires proficiency in Python and familiarity with C++, along with strong software engineering fundamentals and experience with ML accelerators. Candidates should have experience with PyTorch internals and a minimum of 4 years in a relevant software engineering role, preferably with a MS/PhD in Computer Science.

Key Skills:

Python, C++, Triton Kernels, PyTorch, Quantized Inference, Model Compression, Machine Learning Accelerators, vLLM, TRT-LLM, SGLang, Megatron-LM, ModelOpt, Software Engineering, Data Analysis, Numerical Debugging, Large Language Models

Benefits:
  • Equity
  • Benefits
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: Quantized Inference & Pipelines
Senior ML Engineer: Quantized Inference & Pipelines

NVIDIA AI • Redmond (WA)

On-site
USD 140,000 - 200,000
Equity
Benefits
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Senior ML Quantization Engineer for Optical AI Compute
Senior ML Quantization Engineer for Optical AI Compute

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 240,000
Health coverage
HSA contributions
Unlimited PTO
+3
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance