AI Inference Platform Lead – Equity Eligible

NVIDIA

Washington

On-site

USD 184,000 - 357,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA seeks a Manager, Deep Learning Inference Software, to lead a world-class team advancing AI model deployment on NVIDIA GPUs. You will shape software powering cutting-edge AI systems—from large language models to multimodal generative AI—across datacenters and edge devices.

You will drive strategy, roadmap, and execution for inference frameworks, partnering with compiler, libraries, and research teams to deliver end-to-end optimized pipelines and performance tuning for large-scale models.

Qualifications

  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
  • 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management.
  • Strong background in C/C++ software design and development; proficiency in Python is a plus.
  • Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.
  • Proven record of deploying or optimizing deep learning models in production environments.
  • Experience leading teams using Agile or collaborative software development practices.

Responsibilities

  • Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Drive the strategy, roadmap, and execution of NVIDIA's inference frameworks engineering, focusing on Client AI.
  • Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
  • Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).
  • Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA's broader AI and software strategies.
  • Foster a culture of technical excellence, open collaboration, and continuous innovation.

Skills

C/C++ software design
Python
GPU programming
Leadership
Agile practices
Mentoring

Education

MS/PhD in Computer Science or related field

Tools

CUDA Toolkit
Triton
CUTLASS
NCCL

Job description

NVIDIA seeks a Manager, Deep Learning Inference Software, to lead a world-class team advancing AI model deployment on NVIDIA GPUs. You will shape software powering cutting-edge AI systems—from large language models to multimodal generative AI—across datacenters and edge devices.

You will drive strategy, roadmap, and execution for inference frameworks, partnering with compiler, libraries, and research teams to deliver end-to-end optimized pipelines and performance tuning for large-scale models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Software Manager — GPU-Accelerated, Equity
AI Inference Software Manager — GPU-Accelerated, Equity

NVIDIA AI • Santa Clara (UT)

On-site
USD 280,000 - 355,000
Equity
Benefits
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits package
Engineering Manager, GPU AI Inference & Open-Source
Engineering Manager, GPU AI Inference & Open-Source

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits package
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Engineering Manager, AI Inference Platform (Equity)
Senior Engineering Manager, AI Inference Platform (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Engineering Manager, Deep Learning Inference — GPU AI
Engineering Manager, Deep Learning Inference — GPU AI

NVIDIA AI • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity compensation
Benefits
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Head of Deep Learning Inference & OSS Frameworks
Head of Deep Learning Inference & OSS Frameworks

2100 NVIDIA USA • Santa Clara (CA)

Hybrid
USD 224,000 - 432,000