Engineering Manager: GenAI Inference & Deployment at Scale

NVIDIA Gruppe

Santa Clara (CA)

Hybrid

USD 224,000 - 431,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

NVIDIA is seeking an Engineering Manager to lead a team deploying and serving Large Language Models (LLMs) and Vision‑Language Models (VLMs) at scale. You will guide a platform crossing model optimization, inference systems, and distributed infrastructure to deliver high‑performance production deployments.

You will mentor engineers, collaborate with researchers, and shape the roadmap for scalable AI inference, optimizing latency, throughput, and cost across GPU platforms.

Qualifications

  • Master’s or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 8+ years total experience, including 3+ years in management/leadership.
  • Hands-on experience with LLMs/VLMs and modern DL architectures.
  • Expertise in inference optimization techniques: quantization, speculative decoding, batching, KV-cache.

Responsibilities

  • Lead, mentor, and grow a high-performing team building and operating a platform for deploying GenAI models at scale.
  • Drive deployment and optimization of LLMs and VLMs for low-latency, high-throughput inference.
  • Profile and optimize end-to-end DL workloads across model, inference stack, and GPU hardware.
  • Collaborate with research teams to bring new architectures from prototype to production.
  • Set technical strategy and roadmaps for model deployment, inference optimization, and reliability.
  • Establish engineering standards for benchmarking, deployment, and continuous performance optimization.
  • Partner with internal and external teams for seamless GenAI model deployment.

Skills

People management
Leadership
Deep learning
Communication

Education

Master's or PhD in CS/CE

Tools

TensorRT
TensorRT-LLM
vLLM
SGLang

Job description

NVIDIA is seeking an Engineering Manager to lead a team deploying and serving Large Language Models (LLMs) and Vision‑Language Models (VLMs) at scale. You will guide a platform crossing model optimization, inference systems, and distributed infrastructure to deliver high‑performance production deployments.

You will mentor engineers, collaborate with researchers, and shape the roadmap for scalable AI inference, optimizing latency, throughput, and cost across GPU platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, GenAI Inference & Deployment at Scale
Engineering Manager, GenAI Inference & Deployment at Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 320,000 - 420,000
Equity
Benefits
Engineering Manager - Deep Learning Inference on GPUs
Engineering Manager - Deep Learning Inference on GPUs

NVIDIA • Massachusetts

On-site
USD 224,000 - 431,000
Equity compensation
Comprehensive benefits
Engineering Manager, LLM Inference & Deployment at Scale
Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Engineering Manager, LLM Inference & Deployment at Scale
Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 320,000 - 420,000
Equity
Benefits
Engineering Manager, GPU AI Inference at Scale
Engineering Manager, GPU AI Inference at Scale

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits package