TensorRT Performance Engineer for Gen AI Inference

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 124,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Corporation seeks an experienced Deep Learning Software Engineer, TensorRT Performance, to analyze and enhance the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑EdgeLLM, and Torch‑TensorRT.

You will establish benchmarking workflows, optimize model pipelines with quantization and scheduling, and collaborate with teams across AI, automotive, robotics and vision domains to deliver high-performance inference solutions.

Qualifications

  • Bachelor’s, Master’s, PhD, or equivalent experience in CS/CE/AI.
  • At least 2 years of relevant software development experience.
  • Strong C++ and Python programming and software engineering skills.
  • Experience with deep learning frameworks and inference libraries.
  • Experience with performance analysis and optimization.
  • Strong foundation of GPUs and modern DL workloads.

Responsibilities

  • Establish groundbreaking performance benchmarking methodologies and analysis workflows and identify performance issues and opportunities for NVIDIA’s inference ecosystem.
  • Contribute features and code to NVIDIA/OSS inference frameworks including TensorRT, TensorRT‑EdgeLLM, Torch‑TensorRT.
  • Develop new model pipelines with optimized performance including quantization, scheduling, memory management, and distributed inference.
  • Collaborate across generative AI, automotive, robotics, image understanding, and speech understanding teams to develop inference solutions.
  • Scale performance of DL models across architectures and accelerators.

Skills

C++ programming
Python programming
Performance analysis
GPU architecture
DL model workloads
DL frameworks (PyTorch, JAX, Tensor?f)
Inference libraries (TensorRT)
Low-latency/edge AI
Graph compilers (TorchDynamo)
Embedded AI pipelines

Education

Bachelor's/Master's/PhD or equivalent

Tools

CUDA
TileIR
CuTeDSL
cutlass
Triton
TorchDynamo
TorchInductor
vLLM
TorchTensorRT
TensorRT

Job description

NVIDIA Corporation seeks an experienced Deep Learning Software Engineer, TensorRT Performance, to analyze and enhance the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑EdgeLLM, and Torch‑TensorRT.

You will establish benchmarking workflows, optimize model pipelines with quantization and scheduling, and collaborate with teams across AI, automotive, robotics and vision domains to deliver high-performance inference solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
High-Performance AI Inference Engineer (TensorRT)
High-Performance AI Inference Engineer (TensorRT)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 124,000 - 196,000
TensorRT Performance Engineer - Deep Learning
TensorRT Performance Engineer - Deep Learning

NVIDIA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 124,000 - 242,000
Senior Software Engineer, DL Inference (TensorRT)
Senior Software Engineer, DL Inference (TensorRT)

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Senior ML Inference Engineer – TensorRT & LLMs
Senior ML Inference Engineer – TensorRT & LLMs

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits