Stand out for this role — generate a tailored resume and cover letter in about a minute.
Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.
Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA.
Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.
Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA.