TensorRT Performance Engineer

OpenTalent

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is seeking a Deep Learning Software Engineer, TensorRT Performance, to join its research and development team focused on optimizing the inference ecosystem. You will analyze bottlenecks, implement graph compiler algorithms, and improve tensor runtimes across datacenter GPUs and edge accelerators.

In this role, you will collaborate with the deep learning community to integrate TensorRT into OSS frameworks, contribute to Torch-TensorRT and TorchDynamo, and advance state-of-the-art

Qualifications

  • BS/MS/PhD or equivalent in CS/CE/EE/AI.
  • 2+ years of software development experience.
  • Strong C++ and Python programming and software engineering skills.
  • Experience with DL frameworks and inference libraries (TensorRT, PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, FlashInfer).
  • Experience with performance analysis and optimization.

Responsibilities

  • Establish groundbreaking performance benchmarking methodologies and analysis workflows to identify performance issues and opportunities for NVIDIA’s inference ecosystem (TensorRT/TensorRT-EdgeLLM/Torch-TensorRT).
  • Contribute features and code to NVIDIA/OSS inference frameworks, including TensorRT/TensorRT-EdgeLLM/Torch-TensorRT.
  • Develop new model pipelines for NVIDIA’s inference ecosystem with optimized performance, covering quantization, scheduling, memory management, and distributed inference, to set the gold standard for Gen AI performance.
  • Work with cross-collaborative teams inside and outside of NVIDIA across generative AI, automotive, robotics, image understanding, and speech understanding to set directions and develop innovative inference solutions.
  • Scale performance of deep learning models across different architectures and types of NVIDIA accelerators.

Skills

C++
Python
Performance analysis
DL frameworks

Education

Bachelor's degree or higher in CS/CE/EE/AI

Tools

TensorRT
TensorRT-EdgeLLM
Torch-TensorRT
PyTorch
JAX
TensorFlow
ONNX
vLLM
SGLang
FlashInfer
CUDA
TileIR
CuTeDSL
cutlass
Triton

Job description

NVIDIA is seeking a Deep Learning Software Engineer, TensorRT Performance, to join its research and development team focused on optimizing the inference ecosystem. You will analyze bottlenecks, implement graph compiler algorithms, and improve tensor runtimes across datacenter GPUs and edge accelerators.

In this role, you will collaborate with the deep learning community to integrate TensorRT into OSS frameworks, contribute to Torch-TensorRT and TorchDynamo, and advance state-of-the-art

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026

OpenTalent • California (MO)

On-site
USD 150,000 - 210,000
Senior DL Inference Engineer — GPU Performance OpenSource
Senior DL Inference Engineer — GPU Performance OpenSource

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits package
Competitive salary
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
AI Inference Platform Engineer
AI Inference Platform Engineer

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4
Senior AI Performance Engineer
Senior AI Performance Engineer

Brillfy Technology Inc • United States

On-site
USD 150,000 - 210,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Infrastructure Engineer, TensorRT Edge-LLM — Equity
Senior Infrastructure Engineer, TensorRT Edge-LLM — Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 190,000 - 230,000
Equity
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits
GPU Transformer Performance Engineer (Triton/CUDA)
GPU Transformer Performance Engineer (Triton/CUDA)

Luma AI • United States

Remote
USD 180,000 - 280,000