Engineering Manager - Deep Learning Inference on GPUs

NVIDIA

Massachusetts

On-site

USD 224,000 - 431,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity compensation
Comprehensive benefits

Job summary

NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. The role focuses on shaping software powering today’s AI systems — from large language models to multimodal generative AI — all accelerated on NVIDIA GPUs.

The Deep Learning Inference team develops and optimizes open-source frameworks such as SGLang, vLLM, and FlashInfer, enabling scalable, efficient AI deployment.

Qualifications

  • MS/PhD or equivalent in CS/EE or related field.
  • 6+ years in software development.
  • 3+ years in technical leadership or engineering management.
  • Strong background in C/C++ software design.
  • Python proficiency a plus.
  • Hands-on GPU programming (CUDA/Triton/CUTLASS).
  • Experience deploying or optimizing DL models in production.
  • Leadership with Agile or collaborative practices.

Responsibilities

  • Lead and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Define strategy, roadmap, and execution for OSS inference frameworks engineering.
  • Collaborate with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines.
  • Oversee performance tuning and optimization of large-scale models for LLM, multimodal, and generative AI applications.
  • Guide engineers in CUDA, Triton, CUTLASS and multi-GPU communications (NIXL, NCCL, NVSHMEM).
  • Represent the team in roadmap discussions and align with broader NVIDIA strategies.
  • Foster a culture of technical excellence and continuous innovation.

Skills

C/C++ design
Python
CUDA
Triton
CUTLASS
Agile
Production deployment

Education

MS/PhD or equivalent

Tools

SGLang
vLLM
FlashInfer

Job description

NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. The role focuses on shaping software powering today’s AI systems — from large language models to multimodal generative AI — all accelerated on NVIDIA GPUs.

The Deep Learning Inference team develops and optimizes open-source frameworks such as SGLang, vLLM, and FlashInfer, enabling scalable, efficient AI deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, GPU DL Inference & OSS Frameworks
Engineering Manager, GPU DL Inference & OSS Frameworks

NVIDIA • Santa Clara (CA)

Hybrid
USD 224,000 - 431,250
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits package
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager: Deep Learning Inference & GPU
Engineering Manager: Deep Learning Inference & GPU

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Engineering Manager, Deep Learning Inference — GPU AI
Engineering Manager, Deep Learning Inference — GPU AI

NVIDIA AI • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity compensation
Benefits
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Engineering Manager, GPU AI Inference & Open-Source
Engineering Manager, GPU AI Inference & Open-Source

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits package