Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA

Washington (District of Columbia)

On-site

USD 184,000 - 288,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference to design, build, and optimize GPU-accelerated software powering AI applications. You will work on high-performance open-source frameworks, enabling efficient model serving and deployment across datacenters and edge devices.

You will collaborate with the deep learning community to implement the latest techniques for public release in inference frameworks and contribute to NVIDIA libraries such as vLLM, SGLang, and

Qualifications

  • Masters or PhD (or equivalent) in Computer Engineering, Computer Science, EECS, AI.
  • 5+ years of relevant software development experience.
  • Excellent C/C++ programming and software design skills.
  • Python experience is a plus; training/deploying inference experience is a plus.
  • GPU programming with CUDA, OAI Triton or CUTLASS is a plus.

Responsibilities

  • Performance optimization, analysis, and tuning of DL models across domains like LLM, Multimodal, and Generative AI.
  • Scale DL model performance across architectures and NVIDIA accelerators.
  • Contribute features and code to NVIDIA's inference libraries (e.g., vLLM, SGLang) and related tooling.
  • Collaborate with cross-functional teams across frameworks, NVIDIA libraries, and inference optimization.

Skills

C/C++ programming
Software design
Python
CUDA programming
Performance modeling
DL inference

Education

Masters or PhD or equivalent (CS/EECS/AI)

Tools

CUTLASS
OAI Triton
NCCL
CUDA kernels

Job description

NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference to design, build, and optimize GPU-accelerated software powering AI applications. You will work on high-performance open-source frameworks, enabling efficient model serving and deployment across datacenters and edge devices.

You will collaborate with the deep learning community to implement the latest techniques for public release in inference frameworks and contribute to NVIDIA libraries such as vLLM, SGLang, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior DL Inference Engineer — GPU Performance OpenSource
Senior DL Inference Engineer — GPU Performance OpenSource

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits package
Competitive salary
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA AI • Washington

On-site
USD 140,000 - 230,000
Equity
Comprehensive benefits package
Senior DL Inference Engineer — GPU AI, Open Source, Equity
Senior DL Inference Engineer — GPU AI, Open Source, Equity

NVIDIA AI • California (MO)

On-site
USD 150,000 - 240,000
Equity
Health Insurance
Comprehensive Benefits Package
Senior DL Inference Architect — Open-Source + Equity
Senior DL Inference Architect — Open-Source + Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity and benefits
Career growth
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
New Grad GPU Deep Learning Inference Engineer - Equity
New Grad GPU Deep Learning Inference Engineer - Equity

NVIDIA AI • California (MO)

On-site
USD 140,000 - 190,000
Equity
Benefits