Stand out for this role — generate a tailored resume and cover letter in about a minute.
NVIDIA is seeking a Deep Learning Software Engineer, TensorRT Performance, to join its research and development team focused on optimizing the inference ecosystem. You will analyze bottlenecks, implement graph compiler algorithms, and improve tensor runtimes across datacenter GPUs and edge accelerators.
In this role, you will collaborate with the deep learning community to integrate TensorRT into OSS frameworks, contribute to Torch-TensorRT and TorchDynamo, and advance state-of-the-art
NVIDIA is looking for a Deep Learning Software Engineer, TensorRT Performance to join their rapidly growing research and development team for Deep Learning Inference. This role focuses on analyzing and improving the performance of NVIDIA’s inference ecosystem. Companies worldwide leverage NVIDIA GPUs for deep learning, driving breakthroughs in Generative AI, Recommenders, and Vision. The successful candidate will join a team dedicated to building software for performance optimization, deployment, and serving of DL inference solutions, specializing in GPU-accelerated deep learning inference software like TensorRT, DL benchmarking, and performant model deployment solutions.
You will collaborate with the deep learning community to integrate TensorRT into OSS frameworks like TensorRT-EdgeLLM and PyTorch. Key responsibilities include identifying performance opportunities, optimizing state-of-the‑art models across NVIDIA accelerators (from datacenter GPUs to edge SoCs), and implementing graph compiler algorithms, frontend operators, and code generators within NVIDIA’s inference ecosystem. You will also work with various teams on workflow improvements, performance modeling, analysis, kernel development, and inference software development.