Get more replies from employers
Send a job-specific resume in minutes.
Intel is seeking a performance-focused AI Infrastructure Engineer to push LLM inference to its limits on Intel GPUs. You will optimize end-to-end inference pipelines, profile cross-stack bottlenecks, and develop high-performance kernels for attention, MoE, and quantization.
You will also upstream improvements into vLLM, SGLang, and PyTorch while shaping future GPU roadmaps. Work spans from kernel development to open-source contributions, with a hybrid work model and a competitive total
Intel is seeking a performance-focused AI Infrastructure Engineer to push LLM inference to its limits on Intel GPUs. You will optimize end-to-end inference pipelines, profile cross-stack bottlenecks, and develop high-performance kernels for attention, MoE, and quantization.
You will also upstream improvements into vLLM, SGLang, and PyTorch while shaping future GPU roadmaps. Work spans from kernel development to open-source contributions, with a hybrid work model and a competitive total