Turn this role into an interview — a resume and cover letter built around what this employer wants.
NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks.
The role requires strong Python and C++, hands-on GPU profiling (CUPTI/NSYS/NCU), and experience with TRT-LLM, SGLang, or vLLM. Collaboration with compiler, hardware, kernel, and framework teams is essential to deliver
NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks.
The role requires strong Python and C++, hands-on GPU profiling (CUPTI/NSYS/NCU), and experience with TRT-LLM, SGLang, or vLLM. Collaboration with compiler, hardware, kernel, and framework teams is essential to deliver