Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams to optimize throughput and interactivity across Megatron-LM, ModelOpt, and vLLM.
The role requires strong Python skills with familiarity in C++, experience with PyTorch internals, and 4+ years in software engineering.
NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams to optimize throughput and interactivity across Megatron-LM, ModelOpt, and vLLM.
The role requires strong Python skills with familiarity in C++, experience with PyTorch internals, and 4+ years in software engineering.