GPU-Optimized LLM Inference Engineer

Intel Corporation

Folsom (CA)

Hybrid

USD 171,000 - 315,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intel Corporation is seeking a performance-obsessed AI Infrastructure Engineer to push LLM inference to the absolute limits on Intel GPUs, profiling bottlenecks and coding high-performance kernels across the stack.

You will upstream improvements to vLLM, SGLang and PyTorch, collaborate with architecture and compiler teams to shape future GPU roadmaps, and drive performance in state-of-the-art generative AI workloads.

Qualifications

  • Bachelors degree with 4+ years in GPU computing or AI systems
  • Masters degree with 3+ years experience OR PhD
  • Proficient in modern C++ and Python

Responsibilities

  • Own end-to-end optimization pipeline for running LLMs on Intel GPUs.
  • Profile bottlenecks and diagnose cross-stack performance issues.
  • Develop and optimize custom GPU kernels for attention, MoE, quantization.
  • Upstream improvements to vLLM, SGLang, and PyTorch.
  • Collaborate with architecture and compiler teams to shape GPU roadmaps.
  • Contribute to AI infrastructure performance initiatives.

Skills

C++
Python
GPU computing

Education

Bachelor's degree in CS/SE/AI
Master's degree
PhD

Tools

Triton
SYCL
CUDA/CUTLASS
vLLM
SGLang
PyTorch

Job description

Intel Corporation is seeking a performance-obsessed AI Infrastructure Engineer to push LLM inference to the absolute limits on Intel GPUs, profiling bottlenecks and coding high-performance kernels across the stack.

You will upstream improvements to vLLM, SGLang and PyTorch, collaborate with architecture and compiler teams to shape future GPU roadmaps, and drive performance in state-of-the-art generative AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
AI Infra Engineer — Peak LLM on GPUs (Hybrid)
AI Infra Engineer — Peak LLM on GPUs (Hybrid)

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel Corporation • Folsom (CA)

Hybrid
USD 171,000 - 315,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
GPU Inference Engineer — Deep Learning
GPU Inference Engineer — Deep Learning

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1