LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel

Folsom (CA)

Hybrid

USD 171,000 - 315,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock bonuses
Health benefits
Hybrid work model

Job summary

Intel is seeking a performance-focused AI Infrastructure Engineer to push LLM inference to its limits on Intel GPUs. You will optimize end-to-end inference pipelines, profile cross-stack bottlenecks, and develop high-performance kernels for attention, MoE, and quantization.

You will also upstream improvements into vLLM, SGLang, and PyTorch while shaping future GPU roadmaps. Work spans from kernel development to open-source contributions, with a hybrid work model and a competitive total

Qualifications

  • 3+ years of relevant software engineering experience in GPU computing, AI systems, or HPC.
  • Proficiency in modern C++ and Python; comfortable with complex systems-level code.

Responsibilities

  • Drive Inference Performance: Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs.
  • Deep Stack Optimization: Profile, diagnose, and resolve cross-stack performance bottlenecks.
  • Kernel Development and Integration: Design, write, and optimize custom high-performance kernels for attention, MoE, quantization, and operator fusions.
  • Open Source Leadership: Upstream architectural improvements and hardware backends into open-source repos like vLLM, SGLang, PyTorch.
  • Shape the Hardware Roadmap: Apply roofline analysis and profiling to decompose bottlenecks; partner with architecture and compiler teams.

Skills

C++
Python
GPU computing

Education

Bachelor's degree in relevant field
Master's degree (optional)
PhD (optional)

Tools

Triton
SYCL
CUDA/CUTLASS

Job description

Intel is seeking a performance-focused AI Infrastructure Engineer to push LLM inference to its limits on Intel GPUs. You will optimize end-to-end inference pipelines, profile cross-stack bottlenecks, and develop high-performance kernels for attention, MoE, and quantization.

You will also upstream improvements into vLLM, SGLang, and PyTorch while shaping future GPU roadmaps. Work spans from kernel development to open-source contributions, with a hybrid work model and a competitive total

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU-Optimized LLM Inference Engineer
GPU-Optimized LLM Inference Engineer

Intel Corporation • Folsom (CA)

Hybrid
USD 171,000 - 315,000
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
AI Infra Engineer — Peak LLM on GPUs (Hybrid)
AI Infra Engineer — Peak LLM on GPUs (Hybrid)

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior GPU Kernel Optimizer for LLM Inference
Senior GPU Kernel Optimizer for LLM Inference

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000