Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer specializing in Quantized Inference to speed up LLM deployment. You will implement quantized and sparse recipes in inference engines, optimize export pipelines, and build benchmarks for throughput and interactivity across Megatron-LM, vLLM, and related tools.
Key work includes writing Triton kernels, integrating quantize/dequantize paths, and collaborating with cross-team inference groups to push performance.
NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer specializing in Quantized Inference to speed up LLM deployment. You will implement quantized and sparse recipes in inference engines, optimize export pipelines, and build benchmarks for throughput and interactivity across Megatron-LM, vLLM, and related tools.
Key work includes writing Triton kernels, integrating quantize/dequantize paths, and collaborating with cross-team inference groups to push performance.