Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking a Senior Inference Performance Engineer to push the performance limits of large-scale AI inference benchmarks. You will optimize autonomous frameworks used by AI agents to run benchmarks, profile, and tune processes, aiming to maximize throughput per GPU while maintaining model correctness.
You will collaborate with TensorRT-LLM, vLLM, and other teams to translate profiling insights into delivered performance improvements, using Nsight systems and CUDA-based tools to drive
NVIDIA is seeking a Senior Inference Performance Engineer to push the performance limits of large-scale AI inference benchmarks. You will optimize autonomous frameworks used by AI agents to run benchmarks, profile, and tune processes, aiming to maximize throughput per GPU while maintaining model correctness.
You will collaborate with TensorRT-LLM, vLLM, and other teams to translate profiling insights into delivered performance improvements, using Nsight systems and CUDA-based tools to drive