Senior GPU AI Platforms Engineer - Edge LLM Inference

NVIDIA

Durham (NC)

On-site

USD 224,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA’s Local AI team is building the software stack for running large language models and generative AI applications efficiently on NVIDIA edge AI hardware. This role focuses on performance analysis, model validation, and developing inference recipes across multi-node configurations.

The candidate will work with CUDA/C++, Triton, and Python, evaluating new architectures, implementing optimizations, and collaborating with communities and partners to ensure robust model bring-up on NVIDIA GPUs.

Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference
  • Strong Python or C++ programming, software design, and software engineering skills.
  • Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) — you understand how thread blocks, memory hierarchy, and warp execution affect real-world performance
  • Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism
  • Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit
  • Strong analytical skills: ability to form a performance hypothesis, design an experiment, interpret results, and communicate findings clearly

Responsibilities

  • Track and evaluate innovations in leading open-source LLM inference frameworks — identify performance-critical features and algorithmic improvements relevant to NVIDIA edge AI hardware
  • Analyze how new model architectures and inference algorithms map onto NVIDIA GPU architecture — identify mismatch, fallback paths, and optimization opportunities
  • Characterize multi-node inference behavior: collective communication primitives (NCCL/RCCL), topology-aware all-reduce strategies, and parallelism efficiency on edge cluster configurations
  • Produce performance analysis reports mapping theoretical hardware limits to observed inference throughput, latency, and utilization
  • Own the model validation workflow for new model releases: architecture compatibility assessment, inference recipe development, performance characterization, and publication to developer recipe sites
  • Develop and maintain developer-facing inference recipes: keep them accurate as frameworks evolve, automate staleness detection, and build feedback loops from CI results to recipe updates
  • Engage with community and partners on model bring-up questions; serve as the technical point of contact for hardware-specific inference issues related to partner concerns

Skills

Python
C++
CUDA
Triton
Performance analysis
Container tooling

Education

BS/MS/PhD in CS/CE/EE

Tools

Docker
OCI

Job description

NVIDIA’s Local AI team is building the software stack for running large language models and generative AI applications efficiently on NVIDIA edge AI hardware. This role focuses on performance analysis, model validation, and developing inference recipes across multi-node configurations.

The candidate will work with CUDA/C++, Triton, and Python, evaluating new architectures, implementing optimizations, and collaborating with communities and partners to ensure robust model bring-up on NVIDIA GPUs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU AI Platform Engineer — Edge Inference (Equity)
Senior GPU AI Platform Engineer — Edge Inference (Equity)

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior GPU ML Inference Engineer — Edge AI Platforms
Senior GPU ML Inference Engineer — Edge AI Platforms

NVIDIA • Westford (MA)

On-site
USD 224,000 - 432,000
Senior AI Systems Engineer – Edge GPU & Local Inference
Senior AI Systems Engineer – Edge GPU & Local Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Westford (MA)

On-site
USD 224,000 - 432,000
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Austin (TX)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior AI Inference Engineer — GPU & Edge Optimization
Senior AI Inference Engineer — GPU & Edge Optimization

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000