AI Inference Platform Lead – Equity Eligible

NVIDIA

Washington

On-site

USD 184,000 - 357,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA seeks a Manager, Deep Learning Inference Software, to lead a world-class team advancing AI model deployment on NVIDIA GPUs. You will shape software powering cutting-edge AI systems—from large language models to multimodal generative AI—across datacenters and edge devices.

You will drive strategy, roadmap, and execution for inference frameworks, partnering with compiler, libraries, and research teams to deliver end-to-end optimized pipelines and performance tuning for large-scale models.

Qualifications

  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
  • 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.
  • Strong background in C/C++ software design and development; proficiency in Python is a plus.
  • Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.
  • Proven record of deploying or optimizing deep learning models in production environments.
  • Experience leading teams using Agile or collaborative software development practices.

Responsibilities

  • Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Drive the strategy, roadmap, and execution of NVIDIA's inference frameworks engineering, focusing on Client AI.
  • Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
  • Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).
  • Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA's broader AI and software strategies.
  • Foster a culture of technical excellence, open collaboration, and continuous innovation.

Skills

C/C++ software design
Python
GPU programming
Leadership
Agile practices
Mentoring

Education

MS/PhD in Computer Science or related field

Tools

CUDA Toolkit
Triton
CUTLASS
NCCL

Job description

NVIDIA seeks a Manager, Deep Learning Inference Software, to lead a world-class team advancing AI model deployment on NVIDIA GPUs. You will shape software powering cutting-edge AI systems—from large language models to multimodal generative AI—across datacenters and edge devices.

You will drive strategy, roadmap, and execution for inference frameworks, partnering with compiler, libraries, and research teams to deliver end-to-end optimized pipelines and performance tuning for large-scale models.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering Manager, AI Inference & GPU Scaling
Engineering Manager, AI Inference & GPU Scaling

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model
Engineering Manager, GPU AI Inference at Scale
Engineering Manager, GPU AI Inference at Scale

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

NVIDIA • Washington

On-site
USD 224,000 - 431,250
Equity
Benefits package
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 431,250
Equity
Benefits
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 431,250
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager - Deep Learning Inference on GPUs
Engineering Manager - Deep Learning Inference on GPUs

NVIDIA • Massachusetts

On-site
USD 224,000 - 431,000
Equity compensation
Comprehensive benefits
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior Deep Learning Inference Engineer - Equity Eligible
Senior Deep Learning Inference Engineer - Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits