Get more replies from employers
Send a job-specific resume in minutes.
NVIDIA AI in Washington designs, builds, and optimizes GPU-accelerated software for high-performance deep learning inference and model serving.
This role focuses on improving performance for LLM and Generative AI models across NVIDIA accelerators using open-source frameworks, requiring advanced C/C++ skills, GPU programming, and DL optimization expertise.
Design, build, and optimize GPU-accelerated software for high-performance deep learning inference and model serving. Focus on improving performance for LLM and Generative AI models across NVIDIA accelerators using open-source frameworks.
Requires a Master's or PhD in a relevant field and over 5 years of software development experience with strong C/C++ skills. Experience with GPU programming, DL model optimization, and performance profiling is highly desired.
C++, Python, CUDA, Deep Learning Inference, LLM Optimization, Generative AI, OAI Triton, CUTLASS, NCCL, vLLM, SGLang, FlashInfer, Software Design, Performance Profiling, GPU Programming, Agile
Equity, Comprehensive benefits package