Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA’s Local AI team is building the software stack for running large language models and generative AI applications efficiently on NVIDIA edge AI hardware. This role focuses on performance analysis, model validation, and developing inference recipes across multi-node configurations.
The candidate will work with CUDA/C++, Triton, and Python, evaluating new architectures, implementing optimizations, and collaborating with communities and partners to ensure robust model bring-up on NVIDIA GPUs.
NVIDIA’s Local AI team is building the software stack for running large language models and generative AI applications efficiently on NVIDIA edge AI hardware. This role focuses on performance analysis, model validation, and developing inference recipes across multi-node configurations.
The candidate will work with CUDA/C++, Triton, and Python, evaluating new architectures, implementing optimizations, and collaborating with communities and partners to ensure robust model bring-up on NVIDIA GPUs.