Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
NVIDIA AI is seeking an engineer to implement quantized and sparse recipes in inference engines and to manage model export pipelines for correct serialization. You will build benchmarking harnesses and data analysis tools to improve developer productivity through infrastructure and CI improvements.
The role requires strong Python and C++ skills, experience with ML accelerators, and familiarity with PyTorch internals, with 4+ years in software engineering. MS/PhD in CS is preferred.
Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization. Develop benchmarking harnesses, data analysis tools, and improve developer productivity through infrastructure and CI improvements.
Requires proficiency in Python and familiarity with C++, along with strong software engineering fundamentals and experience with ML accelerators. Candidates should have experience with PyTorch internals and a minimum of 4 years in a relevant software engineering role, preferably with a MS/PhD in Computer Science.
Python, C++, Triton Kernels, PyTorch, Quantized Inference, Model Compression, Machine Learning Accelerators, vLLM, TRT-LLM, SGLang, Megatron-LM, ModelOpt, Software Engineering, Data Analysis, Numerical Debugging, Large Language Models