Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 124,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Corporation seeks an experienced Deep Learning Software Engineer, TensorRT Performance, to analyze and enhance the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑EdgeLLM, and Torch‑TensorRT.

You will establish benchmarking workflows, optimize model pipelines with quantization and scheduling, and collaborate with teams across AI, automotive, robotics and vision domains to deliver high-performance inference solutions.

Qualifications

  • Bachelor’s, Master’s, PhD, or equivalent experience in CS/CE/AI.
  • At least 2 years of relevant software development experience.
  • Strong C++ and Python programming and software engineering skills.
  • Experience with deep learning frameworks and inference libraries.
  • Experience with performance analysis and optimization.
  • Strong foundation of GPUs and modern DL workloads.

Responsibilities

  • Establish groundbreaking performance benchmarking methodologies and analysis workflows and identify performance issues and opportunities for NVIDIA’s inference ecosystem.
  • Contribute features and code to NVIDIA/OSS inference frameworks including TensorRT, TensorRT‑EdgeLLM, Torch‑TensorRT.
  • Develop new model pipelines with optimized performance including quantization, scheduling, memory management, and distributed inference.
  • Collaborate across generative AI, automotive, robotics, image understanding, and speech understanding teams to develop inference solutions.
  • Scale performance of DL models across architectures and accelerators.

Skills

C++ programming
Python programming
Performance analysis
GPU architecture
DL model workloads
DL frameworks (PyTorch, JAX, Tensor?f)
Inference libraries (TensorRT)
Low-latency/edge AI
Graph compilers (TorchDynamo)
Embedded AI pipelines

Education

Bachelor's/Master's/PhD or equivalent

Tools

CUDA
TileIR
CuTeDSL
cutlass
Triton
TorchDynamo
TorchInductor
vLLM
TorchTensorRT
TensorRT

Job description

Job Overview

NVIDIA is seeking an experienced Deep Learning Software Engineer, TensorRT Performance to analyze and improve the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑EdgeLLM and Torch‑TensorRT.

Responsibilities
  • Establish groundbreaking performance benchmarking methodologies and analysis workflows and identify performance issues and opportunities for NVIDIA’s inference ecosystem (e.g. TensorRT, TensorRT‑EdgeLLM, Torch‑TensorRT).
  • Contribute features and code to NVIDIA/OSS inference frameworks including but not limited to TensorRT, TensorRT‑EdgeLLM, Torch‑TensorRT.
  • Develop new model pipelines for NVIDIA’s inference ecosystem with optimized performance including quantization, scheduling, memory management, and distributed inference to set the gold standard for Gen AI performance.
  • Work with cross‑collaborative teams inside and outside of NVIDIA across generative AI, automotive, robotics, image understanding, and speech understanding to set directions and develop innovative inference solutions.
  • Scale performance of deep learning models across different architectures and types of NVIDIA accelerators.
Qualifications
  • Bachelor’s, Master’s, PhD, or equivalent experience in Computer Science, Computer Engineering, EECS, or AI.
  • At least 2 years of relevant software development experience.
  • Strong C++ and Python programming and software engineering skills.
  • Experience with deep learning frameworks (PyTorch, JAX, TensorFlow, ONNX) and inference libraries (TensorRT, TensorRT‑LLM, vLLM, SGLang, FlashInfer).
  • Experience with performance analysis and performance optimization.
  • Strong foundation and architectural knowledge of GPUs.
  • Deep understanding of modern deep learning models and workloads (Transformers, Recommenders, ASR, TTS, Visual Understanding).
  • Proficiency in at least one deep learning programming domain‑specific language (CUDA, TileIR, CuTeDSL, cutlass, Triton).
  • Prior contributions to major LLM inference frameworks (e.g. vLLM) or experience with graph compilers in deep learning inference (e.g. TorchDynamo, TorchInductor).
  • Prior experience optimizing performance for low‑latency, resource‑constrained systems or embedded AI pipelines (e.g. Jetson systems, other edge AI accelerators).
Compensation

Base salary ranges: Level2: $124,000–$195,500 USD; Level3: $152,000–$241,500 USD. Eligible for equity and benefits.

Application Window

Applications will be accepted until at least July24,2026.

EEO Statement

NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026

NVIDIA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
TensorRT Performance Engineer - Deep Learning
TensorRT Performance Engineer - Deep Learning

NVIDIA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Deep Learning Software Engineer, Inference - New College Grad 2026
Deep Learning Software Engineer, Inference - New College Grad 2026

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Software Engineer, TensorRT Specialized Platforms - New College Grad 2025
Software Engineer, TensorRT Specialized Platforms - New College Grad 2025

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits