Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits on large-scale AI inference benchmarks. You will optimize AI model execution and profiling workflows, enabling autonomous optimization across TensorRT-LLM, vLLM, and disaggregated serving architectures.
The role emphasizes measurable gains, reproducible experiments, and close collaboration with GPU and software teams. Responsibilities include distilling performance methods, benchmarking with Nsight tools, and
NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits on large-scale AI inference benchmarks. You will optimize AI model execution and profiling workflows, enabling autonomous optimization across TensorRT-LLM, vLLM, and disaggregated serving architectures.
The role emphasizes measurable gains, reproducible experiments, and close collaboration with GPU and software teams. Responsibilities include distilling performance methods, benchmarking with Nsight tools, and