Engineering Manager, GenAI Inference & Deployment at Scale

NVIDIA

Santa Clara (CA)

Hybrid

USD 320,000 - 420,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA seeks an Engineering Manager to lead a team building and operating a platform that deploys GenAI models at scale on NVIDIA GPUs. You will collaborate with researchers, software engineers, and hardware specialists to turn research into reliable, production deployments.

Ideal candidates have 8+ years of experience including 3+ in management, a Master’s/PhD, and hands-on expertise in LLMs/VLMs and inference optimization.

Qualifications

  • 8+ years total experience with at least 3 years in management/leadership.
  • Strong hands-on experience with LLMs and/or VLMs and understanding of modern DL architectures.
  • Deep expertise in inference optimization techniques and GPU-accelerated deployment.
  • Experience with disaggregated inference/serving, distributed inference, and multi-node deployments.

Responsibilities

  • Lead, mentor, and grow a high-performing team building and operating a platform for deploying GenAI models at scale.
  • Drive the deployment and optimization of LLMs and VLMs for low-latency, high-throughput, and cost-efficient inference.
  • Analyze, profile, and optimize end-to-end deep learning workloads across the model, inference stack, distributed systems, and GPU hardware.
  • Work closely with research teams and model developers to bring new architectures and models from prototype to production.
  • Drive technical strategy and roadmap for model deployment, inference optimization, scalability, reliability, and performance.
  • Establish engineering standards for benchmarking, profiling, production deployment, and continuous performance optimization.
  • Collaborate with internal and external partners to enable seamless deployment of rapidly evolving GenAI models.

Skills

Leadership
Deep learning
Distributed inference
Model optimization

Education

Master's or PhD in CS/CE or related field

Tools

TensorRT
TensorRT-LLM
vLLM
SGLang

Job description

NVIDIA seeks an Engineering Manager to lead a team building and operating a platform that deploys GenAI models at scale on NVIDIA GPUs. You will collaborate with researchers, software engineers, and hardware specialists to turn research into reliable, production deployments.

Ideal candidates have 8+ years of experience including 3+ in management, a Master’s/PhD, and hands-on expertise in LLMs/VLMs and inference optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager: GenAI Inference & Deployment at Scale
Engineering Manager: GenAI Inference & Deployment at Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Engineering Manager, GPU AI Inference at Scale
Engineering Manager, GPU AI Inference at Scale

NVIDIA • New York (NY)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager - Deep Learning Inference on GPUs
Engineering Manager - Deep Learning Inference on GPUs

NVIDIA • Massachusetts

On-site
USD 224,000 - 431,000
Equity compensation
Comprehensive benefits
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Engineering Manager, LLM Inference & Deployment at Scale
Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA • Santa Clara (CA)

Hybrid
USD 320,000 - 420,000
Equity
Benefits
Engineering Manager, LLM Inference & Deployment at Scale
Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 224,000 - 431,000
Equity
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, Deep Learning Algorithms & Inference
Engineering Manager, Deep Learning Algorithms & Inference

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits