Edge AI LLM Inference Architect (C++, TensorRT)

NVIDIA AI

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health Insurance

Job summary

NVIDIA AI in Santa Clara is seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending TensorRT with autoregressive model serving capabilities and requires collaboration across CUDA, kernel libraries, compilers, and robotics teams to deliver high-performance, production-ready solutions.

The candidate should hold a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep

Qualifications

  • Candidates should have a BS, MS, PhD, or equivalent experience in a relevant field and at least 4 years of software development experience.
  • A deep understanding of transformer models and proficiency in modern C++ are essential.

Responsibilities

  • Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities.
  • Collaborate with teams across CUDA, kernel libraries, compilers, and robotics to deliver high-performance, production-ready solutions.

Skills

C++
TensorRT
LLM
VLM
GEMM
CUDA
Attention
MoE
KV Cache Management
Speculative Decoding
Quantization
Tensor Parallelism
Memory-Efficient Scheduling
Compiler Infrastructure
Robotics
Embedded AI

Education

BS/MS/PhD or equivalent

Job description

NVIDIA AI in Santa Clara is seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending TensorRT with autoregressive model serving capabilities and requires collaboration across CUDA, kernel libraries, compilers, and robotics teams to deliver high-performance, production-ready solutions.

The candidate should hold a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer – TensorRT & LLMs
Senior ML Inference Engineer – TensorRT & LLMs

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer – TensorRT Edge-LLM
Senior Software Engineer – TensorRT Edge-LLM

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
Health Insurance
Senior Edge AI Engineer: LLM Inference (TensorRT)
Senior Edge AI Engineer: LLM Inference (TensorRT)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Hybrid work model
Senior DL Inference Engineer – Automotive, C++, TensorRT Equity
Senior DL Inference Engineer – Automotive, C++, TensorRT Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity
Senior ML Inference Engineer - TensorRT/CUDA
Senior ML Inference Engineer - TensorRT/CUDA

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
High-Performance AI Inference Engineer (TensorRT)
High-Performance AI Inference Engineer (TensorRT)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior AI-Native Systems Software Engineer, TensorRT
Senior AI-Native Systems Software Engineer, TensorRT

NVIDIA AI • Santa Clara (CA)

On-site
USD 150,000 - 230,000
Equity
Benefits
AI-Native Systems Engineer, TensorRT — Hybrid & Equity
AI-Native Systems Engineer, TensorRT — Hybrid & Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Hybrid work model
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity
Senior DL Inference Engineer (TensorRT) - Hybrid + Equity

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits