Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026

OpenTalent

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA is seeking a Deep Learning Software Engineer, TensorRT Performance, to join its research and development team focused on optimizing the inference ecosystem. You will analyze bottlenecks, implement graph compiler algorithms, and improve tensor runtimes across datacenter GPUs and edge accelerators.

In this role, you will collaborate with the deep learning community to integrate TensorRT into OSS frameworks, contribute to Torch-TensorRT and TorchDynamo, and advance state-of-the-art

Qualifications

  • BS/MS/PhD or equivalent in CS/CE/EE/AI.
  • 2+ years of software development experience.
  • Strong C++ and Python programming and software engineering skills.
  • Experience with DL frameworks and inference libraries (TensorRT, PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, FlashInfer).
  • Experience with performance analysis and optimization.

Responsibilities

  • Establish groundbreaking performance benchmarking methodologies and analysis workflows to identify performance issues and opportunities for NVIDIA’s inference ecosystem (TensorRT/TensorRT-EdgeLLM/Torch-TensorRT).
  • Contribute features and code to NVIDIA/OSS inference frameworks, including TensorRT/TensorRT-EdgeLLM/Torch-TensorRT.
  • Develop new model pipelines for NVIDIA’s inference ecosystem with optimized performance, covering quantization, scheduling, memory management, and distributed inference, to set the gold standard for Gen AI performance.
  • Work with cross-collaborative teams inside and outside of NVIDIA across generative AI, automotive, robotics, image understanding, and speech understanding to set directions and develop innovative inference solutions.
  • Scale performance of deep learning models across different architectures and types of NVIDIA accelerators.

Skills

C++
Python
Performance analysis
DL frameworks

Education

Bachelor's degree or higher in CS/CE/EE/AI

Tools

TensorRT
TensorRT-EdgeLLM
Torch-TensorRT
PyTorch
JAX
TensorFlow
ONNX
vLLM
SGLang
FlashInfer
CUDA
TileIR
CuTeDSL
cutlass
Triton

Job description

About the Role

NVIDIA is looking for a Deep Learning Software Engineer, TensorRT Performance to join their rapidly growing research and development team for Deep Learning Inference. This role focuses on analyzing and improving the performance of NVIDIA’s inference ecosystem. Companies worldwide leverage NVIDIA GPUs for deep learning, driving breakthroughs in Generative AI, Recommenders, and Vision. The successful candidate will join a team dedicated to building software for performance optimization, deployment, and serving of DL inference solutions, specializing in GPU-accelerated deep learning inference software like TensorRT, DL benchmarking, and performant model deployment solutions.

You will collaborate with the deep learning community to integrate TensorRT into OSS frameworks like TensorRT-EdgeLLM and PyTorch. Key responsibilities include identifying performance opportunities, optimizing state-of-the‑art models across NVIDIA accelerators (from datacenter GPUs to edge SoCs), and implementing graph compiler algorithms, frontend operators, and code generators within NVIDIA’s inference ecosystem. You will also work with various teams on workflow improvements, performance modeling, analysis, kernel development, and inference software development.

What you’ll be doing:
  • Establish groundbreaking performance benchmarking methodologies and analysis workflows to identify performance issues and opportunities for NVIDIA’s inference ecosystem (e.g., TensorRT/TensorRT-EdgeLLM/Torch-TensorRT).
  • Contribute features and code to NVIDIA/OSS inference frameworks, including but not limited to TensorRT/TensorRT-EdgeLLM/Torch-TensorRT.
  • Develop new model pipelines for NVIDIA’s inference ecosystem with optimized performance, covering areas like quantization, scheduling, memory management, and distributed inference, to set the gold standard for Gen AI performance.
  • Work with cross-collaborative teams inside and outside of NVIDIA across generative AI, automotive, robotics, image understanding, and speech understanding to set directions and develop innovative inference solutions.
  • Scale performance of deep learning models across different architectures and types of NVIDIA accelerators.
What we need to see:
  • Bachelors, Masters, PhD, or equivalent experience in relevant fields (Computer Science, Computer Engineering, EECS, AI).
  • 2+ years of relevant software development experience.
  • Strong C++ and Python programming and software engineering skills.
  • Experience with DL frameworks (e.g., PyTorch, JAX, TensorFlow, ONNX) and inference libraries (e.g., TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer).
  • Experience with performance analysis and performance optimization.
Ways to stand out from the crowd:
  • Strong foundation and architectural knowledge of GPUs.
  • Deep understanding of modern deep learning models and workloads (e.g., Transformers, Recommenders, ASR, TTS, Visual Understanding).
  • Proficiency in one of the deep learning programming domain‑specific languages (e.g., CUDA/TileIR/CuTeDSL/cutlass/Triton).
  • Prior contributions to major LLM inference frameworks (e.g., vLLM) or prior experience with graph compilers in deep learning inference (e.g., TorchDynamo/TorchInductor).
  • Prior experience optimizing performance for low‑latency, resource‑constrained systems or embedded AI pipelines (e.g., Jetson systems or other edge AI accelerators).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA AI • Washington

On-site
USD 140,000 - 230,000
Equity
Comprehensive benefits package
TensorRT Performance Engineer
TensorRT Performance Engineer

OpenTalent • California (MO)

On-site
USD 150,000 - 210,000
Senior Performance Engineer - Deep Learning
Senior Performance Engineer - Deep Learning

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior Deep Learning Software Engineer, Inference
Senior Deep Learning Software Engineer, Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Deep Learning Architect, LLM Inference
Senior Deep Learning Architect, LLM Inference

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Deep Learning Software Engineer, Inference - New College Grad 2026
Deep Learning Software Engineer, Inference - New College Grad 2026

NVIDIA AI • California (MO)

On-site
USD 140,000 - 190,000
Equity
Benefits