Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA seeks a Senior Software Engineer – AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-end analysis, build performance models, and optimize latency, throughput, and energy efficiency across models, serving software, and distributed runtimes.
The role spans profiling, tuning, and developing high-performance kernels with CUDA, Triton, and CUTLASS while collaborating with model, kernel, networking, and GPU architecture
NVIDIA seeks a Senior Software Engineer – AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-end analysis, build performance models, and optimize latency, throughput, and energy efficiency across models, serving software, and distributed runtimes.
The role spans profiling, tuning, and developing high-performance kernels with CUDA, Triton, and CUTLASS while collaborating with model, kernel, networking, and GPU architecture