Engineering Manager, Deep Learning Inference

Jobtailor

California (MO)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a senior engineering leader to guide an elite team focused on deep learning inference and GPU-accelerated software. You will shape strategy, drive execution of inference frameworks for Client AI, and collaborate with compiler, libraries, and research teams to optimize pipelines across accelerators.

You will mentor engineers, champion CUDA/Triton/CUTLASS adoption, and push performance optimization for large-scale models. Expect a roles-focused, impact-driven environment.

Qualifications

  • MS/PhD or equivalent in CS/EE or related field.
  • 6+ years of software development experience.
  • 3+ years in technical leadership or engineering management.
  • Strong background in C/C++ software design and development.
  • Proficiency in Python is a plus.
  • Hands-on GPU programming with CUDA/Triton/CUTLASS.
  • Experience with performance optimization in production.
  • Leading teams with Agile or collaborative practices.
  • Open-source contributions to DL or inference frameworks advantageous.

Responsibilities

  • Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Drive strategy, roadmap, and execution of inference frameworks engineering, focusing on Client AI.
  • Partner with compiler, libraries, and research teams to deliver optimized inference pipelines across NVIDIA accelerators.
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM and multimodal AI.
  • Guide engineers in adopting CUDA, Triton, and multi-GPU communication technologies.
  • Represent the team in roadmap and planning discussions and foster technical excellence.

Skills

Technical Leadership
Deep Learning Model Optimization
CUDA Programming
Performance Tuning
Agile Software Development
Team Leadership

Education

MS/PhD in CS/EE or related field

Tools

CUDA
Triton
CUTLASS
NIXL
NCCL
NVSHMEM

Job description

  • Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software
  • Drive the strategy, roadmap, and execution of NVIDIA’s inference frameworks engineering, focusing on Client AI
  • Partner with internal compiler, libraries, and research teams to deliver optimized inference pipelines across NVIDIA accelerators
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications
  • Guide engineers in adopting CUDA, Triton, CUTLASS, and multi-GPU communication technologies
  • Represent the team in roadmap and planning discussions
  • Foster technical excellence, collaboration, and continuous innovation
Requirements
  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field
  • 6+ overall years of software development experience
  • 3+ years in technical leadership or engineering management
  • Strong background in C/C++ software design and development
  • Proficiency in Python is a plus
  • Hands-on experience with GPU programming using CUDA, Triton, and CUTLASS
  • Experience with performance optimization
  • Proven record of deploying or optimizing deep learning models in production environments
  • Experience leading teams using Agile or collaborative software development practices
  • Significant open-source contributions to deep learning or inference frameworks are advantageous
  • Deep understanding of NIXL, NCCL, NVSHMEM, and distributed inference architectures is advantageous
  • Expertise in performance modeling, profiling, and system-level optimization is advantageous
  • Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects is advantageous
  • Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering are advantageous
Core Competencies

Demonstrates expertise in leading and mentoring engineering teams focused on deep learning inference and GPU-accelerated software, with a strong emphasis on performance optimization and deployment of large-scale models. Proficient in CUDA, Triton, and CUTLASS, with a solid foundation in C/C++ and Python programming.

Highest-signal resume keywords
  • Technical Leadership
  • Deep Learning Model Optimization
  • CUDA Programming
  • Performance Tuning
  • Agile Software Development
ATS Optimization Keywords
Hard Skills
  • C/C++ Software Design
  • Python Programming
  • GPU Programming
  • Performance Optimization
  • Deep Learning Frameworks
  • Inference Frameworks
  • Model Deployment
  • System-Level Optimization
  • Performance Profiling
  • Architectural Decision-Making
Soft Skills
  • Mentoring
  • Collaboration
  • Innovation
  • Communication
  • Team Leadership
Industry Keywords
  • Deep Learning
  • GPU Acceleration
  • Inference Pipelines
  • Large-Scale Models
  • Multimodal AI
  • Generative AI
  • Open-Source Contributions
Tools & Technologies
  • CUDA
  • Triton
  • CUTLASS
  • NIXL
  • NCCL
  • NVSHMEM
  • Agile Methodologies
  • Distributed Inference Architectures
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA AI • Santa Clara (UT)

On-site
USD 280,000 - 355,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Massachusetts

On-site
USD 224,000 - 431,000
Equity compensation
Comprehensive benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA AI • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity compensation
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits package
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Socket.dev • Santa Clara (UT)

On-site
USD 224,000 - 357,000
Equity
Benefits package