LLM AI Inference Performance Engineer

Intel

California (MO)

Hybrid

USD 171,000 - 315,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intel is seeking an AI Infrastructure Engineer to push LLM inference on next‑generation GPUs. You will profile stacks, write high‑performance kernels, and upstream optimizations to vLLM, SGLang, and PyTorch, shaping Intel hardware for GenAI workloads.

This hybrid role is based in the United States with primary location Santa Clara, CA, and includes multiple US sites. Strong C++/Python and HPC experience are required, with 4+ years in GPU computing or AI systems.

Qualifications

  • 4+ years in GPU computing or HPC.
  • Proficiency in C++ and Python.
  • Bachelor's degree or higher in CS/SE/AI.

Responsibilities

  • Profile end-to-end inference pipeline for running state-of-the-art LLMs on Intel GPUs.
  • Profile, diagnose, and resolve cross-stack performance bottlenecks.
  • Design, write, and optimize custom high-performance kernels for attention, MoE, quantization, and operator fusions.
  • Upstream architectural improvements to vLLM, SGLang, and PyTorch; bridge hardware teams with the open-source community.
  • Apply roofline analysis and profiling to shape future Intel GPU roadmaps.
  • Show passion about AI infrastructure and performance optimization.

Skills

C++
Python
GPU computing
HPC

Education

Bachelor's degree in CS/SE/AI
Masters or PhD (advantage)

Tools

vLLM
SGLang
PyTorch
CUDA/CUTLASS

Job description

Intel is seeking an AI Infrastructure Engineer to push LLM inference on next‑generation GPUs. You will profile stacks, write high‑performance kernels, and upstream optimizations to vLLM, SGLang, and PyTorch, shaping Intel hardware for GenAI workloads.

This hybrid role is based in the United States with primary location Santa Clara, CA, and includes multiple US sites. Strong C++/Python and HPC experience are required, with 4+ years in GPU computing or AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel Corporation • Folsom (CA)

Hybrid
USD 171,000 - 315,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation