Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance LLM and VLM inference on GPU-accelerated systems. You will lead end-to-end analysis, define workloads, and push latency and throughput improvements across models, serving software, and distributed runtimes.
The role requires hands-on coding in Python/C++/Rust, deep GPU architecture knowledge, and experience profiling with Nsight tools.
NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance LLM and VLM inference on GPU-accelerated systems. You will lead end-to-end analysis, define workloads, and push latency and throughput improvements across models, serving software, and distributed runtimes.
The role requires hands-on coding in Python/C++/Rust, deep GPU architecture knowledge, and experience profiling with Nsight tools.