Senior GPU AI Platforms Engineer - Edge LLM Inference

NVIDIA

Durham (NC)

In loco

USD 224.000 - 356.500

Tempo pieno

14 giorni+
Generatore di candidature

Non inviare un curriculum generico — genera un curriculum e una lettera di presentazione personalizzati per questo specifico impiego.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Equity
Benefits

Descrizione del lavoro

NVIDIA’s Local AI team is building the software stack for running large language models and generative AI applications efficiently on NVIDIA edge AI hardware. This role focuses on performance analysis, model validation, and developing inference recipes across multi-node configurations.

The candidate will work with CUDA/C++, Triton, and Python, evaluating new architectures, implementing optimizations, and collaborating with communities and partners to ensure robust model bring-up on NVIDIA GPUs.

Competenze

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference
  • Strong Python or C++ programming, software design, and software engineering skills.
  • Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) — you understand how thread blocks, memory hierarchy, and warp execution affect real-world performance
  • Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism
  • Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit
  • Strong analytical skills: ability to form a performance hypothesis, design an experiment, interpret results, and communicate findings clearly

Mansioni

  • Track and evaluate innovations in leading open-source LLM inference frameworks — identify performance-critical features and algorithmic improvements relevant to NVIDIA edge AI hardware
  • Analyze how new model architectures and inference algorithms map onto NVIDIA GPU architecture — identify mismatch, fallback paths, and optimization opportunities
  • Characterize multi-node inference behavior: collective communication primitives (NCCL/RCCL), topology-aware all-reduce strategies, and parallelism efficiency on edge cluster configurations
  • Produce performance analysis reports mapping theoretical hardware limits to observed inference throughput, latency, and utilization
  • Own the model validation workflow for new model releases: architecture compatibility assessment, inference recipe development, performance characterization, and publication to developer recipe sites
  • Develop and maintain developer-facing inference recipes: keep them accurate as frameworks evolve, automate staleness detection, and build feedback loops from CI results to recipe updates
  • Engage with community and partners on model bring-up questions; serve as the technical point of contact for hardware-specific inference issues related to partner concerns

Conoscenze

Python
C++
CUDA
Triton
Performance analysis
Container tooling

Formazione

BS/MS/PhD in CS/CE/EE

Strumenti

Docker
OCI

Descrizione del lavoro

NVIDIA’s Local AI team is building the software stack for running large language models and generative AI applications efficiently on NVIDIA edge AI hardware. This role focuses on performance analysis, model validation, and developing inference recipes across multi-node configurations.

The candidate will work with CUDA/C++, Triton, and Python, evaluating new architectures, implementing optimizations, and collaborating with communities and partners to ensure robust model bring-up on NVIDIA GPUs.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior Edge AI GPU & LLM Inference Engineer
Senior Edge AI GPU & LLM Inference Engineer

NVIDIA Corporation • Santa Clara (CA)

In loco
USD 224.000 - 431.000
Senior GPU ML Inference Engineer — Edge AI Platforms
Senior GPU ML Inference Engineer — Edge AI Platforms

NVIDIA • Westford (MA)

In loco
USD 224.000 - 431.250
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

In loco
USD 184.000 - 357.000
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Austin (TX)

In loco
USD 224.000 - 431.250
Equity
Benefits
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Durham (NC)

In loco
USD 224.000 - 356.500
Equity
Benefits
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA • Westford (MA)

In loco
USD 224.000 - 431.250
Senior Software Engineer - GPU Local AI Platforms
Senior Software Engineer - GPU Local AI Platforms

NVIDIA Corporation • Santa Clara (CA)

In loco
USD 224.000 - 431.000
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

In loco
USD 152.000 - 288.000
Senior DL Inference Engineer - GPU & LLM Performance
Senior DL Inference Engineer - GPU & LLM Performance

NVIDIA Corporation • Santa Clara (CA)

In loco
USD 184.000 - 357.000
Equity
Benefits
Senior AI Compiler Engineer - MLIR, GPU Inference, Equity
Senior AI Compiler Engineer - MLIR, GPU Inference, Equity

NVIDIA • Town of Texas (WI)

In loco
USD 152.000 - 288.000
Equity
Benefits package