Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA AI in Santa Clara is seeking a highly capable software engineer to advance an advanced inference framework using modern C++. The role focuses on extending TensorRT with autoregressive model serving capabilities and requires collaboration across CUDA, kernel libraries, compilers, and robotics teams to deliver high-performance, production-ready solutions.
The candidate should hold a BS/MS/PhD (or equivalent) and have at least four years of software development experience with a deep
Develop and evolve a state-of-the-art inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities. Collaborate with teams across CUDA, kernel libraries, compilers, and robotics to deliver high-performance, production-ready solutions.
Candidates should have a BS, MS, PhD, or equivalent experience in a relevant field and at least 4 years of software development experience. A deep understanding of transformer models and proficiency in modern C++ are essential.
C++, TensorRT, LLM, VLM, GEMM, CUDA, Attention, MoE, KV Cache Management, Speculative Decoding, Quantization, Tensor Parallelism, Memory-Efficient Scheduling, Compiler Infrastructure, Robotics, Embedded AI
Equity, Health Insurance