LLM Inference Performance Engineer

Intel

Santa Clara (CA)

Hybrid

USD 171,000 - 315,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intel is seeking an AI Infrastructure Engineer in Santa Clara to push LLM inference on Intel GPUs, optimizing the end-to-end stack from profiling to kernel development. You will upstream improvements to open-source platforms like vLLM, SGLang, and PyTorch and collaborate with architecture teams to shape future GPU roadmaps.

The role requires strong C++/Python skills and 4+ years in GPU computing or HPC, with a passion for performance and scale-out inference across multi-node setups.

Qualifications

  • Bachelor's degree in CS/SE/AI with 4+ years experience or higher degrees with corresponding years.
  • 3+ years in GPU computing, AI systems or HPC.
  • Proficiency in C++ and Python; able to read/modify complex systems code.

Responsibilities

  • Own end-to-end optimization for running state-of-the-art LLMs on Intel GPUs.
  • Profile, diagnose, and resolve cross-stack performance bottlenecks.
  • Design and optimize custom high-performance kernels for attention, MoE, quantization.
  • Upstream improvements to vLLM, SGLang, PyTorch; liaise with hardware and open-source communities.
  • Shape hardware roadmap using roofline analysis and real workload data.
  • Show passion for AI infrastructure and performance optimization.

Skills

C++
Python
GPU computing
Profiling
Kernel development

Education

Bachelors degree in CS/SE/AI

Tools

Triton
CUDA
PyTorch
vLLM
SGLang

Job description

Intel is seeking an AI Infrastructure Engineer in Santa Clara to push LLM inference on Intel GPUs, optimizing the end-to-end stack from profiling to kernel development. You will upstream improvements to open-source platforms like vLLM, SGLang, and PyTorch and collaborate with architecture teams to shape future GPU roadmaps.

The role requires strong C++/Python skills and 4+ years in GPU computing or HPC, with a passion for performance and scale-out inference across multi-node setups.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1