GPU Inference Engineer — Deep Learning

2100 NVIDIA USA

California (MO)

On-site

USD 124,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA seeks a Software Engineer specializing in Deep Learning Inference to design, build and optimize GPU-accelerated software powering AI applications. You will contribute to high-performance inference frameworks, including vLLM, SGLang, and related NVIDIA libraries, across datacenter and edge hardware.

You’ll work with the DL community to implement the latest algorithms for public release, focusing on performance improvements for state-of-the-art LLMs and Generative AI on NVIDIA accelerators.

Qualifications

  • Pursuing or recently completed a MS or PhD in Computer Engineering, Computer Science, EECS, AI or related field.
  • Strong software development experience.
  • Excellent C/C++ programming and software design skills; Python is a plus.
  • Experience with training, deploying or optimizing DL model inference is advantageous.

Responsibilities

  • Performance optimization, analysis, and tuning of DL models across domains like LLMs and Generative AI.
  • Scale DL model performance across architectures and NVIDIA accelerators.
  • Contribute features to inference libraries and open-source frameworks.
  • Collaborate with cross-functional teams across frameworks and NVIDIA libraries.

Skills

C/C++
Software design
Python
Agile methodologies

Education

MS or PhD in CS/EECS/AI

Tools

CUDA
CUTLASS
OAI Triton
NCCL
vLLM
SGLang

Job description

NVIDIA seeks a Software Engineer specializing in Deep Learning Inference to design, build and optimize GPU-accelerated software powering AI applications. You will contribute to high-performance inference frameworks, including vLLM, SGLang, and related NVIDIA libraries, across datacenter and edge hardware.

You’ll work with the DL community to implement the latest algorithms for public release, focusing on performance improvements for state-of-the-art LLMs and Generative AI on NVIDIA accelerators.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager - Deep Learning Inference on GPUs
Engineering Manager - Deep Learning Inference on GPUs

NVIDIA • Massachusetts

On-site
USD 224,000 - 431,000
Equity compensation
Comprehensive benefits
Engineering Manager, Deep Learning Inference — GPU AI
Engineering Manager, Deep Learning Inference — GPU AI

NVIDIA AI • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity compensation
Benefits
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA AI • California (MO)

On-site
USD 184,000 - 288,000
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits