Senior AI Inference & Kernel Engineer

Intel

Austin (TX)

Hybrid

USD 189,000 - 315,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock bonuses
Health benefits
Vacation

Job summary

Intel is seeking an AI Infrastructure Engineer to push LLM inference on Intel GPUs to new limits in Santa Clara and across multiple US locations. You will profile bottlenecks, write high-performance GPU kernels, and upstream improvements to open-source projects like vLLM and PyTorch.

You will work end-to-end across the stack, shape hardware roadmaps with roofline analysis, and collaborate with architecture and compiler teams to optimize performance for state-of-the-art generative AI workloads.

Qualifications

  • Experience in GPU computing and AI systems development.
  • Strong proficiency in C++ and Python; reading/modifying complex code.
  • Higher degree or equivalent experience is preferred.

Responsibilities

  • Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs.
  • Profile, diagnose, and resolve cross-stack performance bottlenecks.
  • Design and optimize custom high-performance kernels for attention, MoE, quantization, and fusions.
  • Upstream architectural improvements into open-source projects like vLLM, SGLang, PyTorch.
  • Collaborate with architecture and compiler teams to shape GPU roadmaps based on GenAI workloads.
  • Demonstrate passion for AI infrastructure and performance optimization.

Skills

GPU computing
C++
Python
HPC

Education

Bachelor's degree in Computer Science or related field

Tools

Triton
CUDA/CUTLASS
SYCL

Job description

Intel is seeking an AI Infrastructure Engineer to push LLM inference on Intel GPUs to new limits in Santa Clara and across multiple US locations. You will profile bottlenecks, write high-performance GPU kernels, and upstream improvements to open-source projects like vLLM and PyTorch.

You will work end-to-end across the stack, shape hardware roadmaps with roofline analysis, and collaborate with architecture and compiler teams to optimize performance for state-of-the-art generative AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
GPU-Optimized LLM Inference Engineer
GPU-Optimized LLM Inference Engineer

Intel Corporation • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
AI Infra Engineer — Peak LLM on GPUs (Hybrid)
AI Infra Engineer — Peak LLM on GPUs (Hybrid)

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel Corporation • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
AI Infrastructure Engineer
AI Infrastructure Engineer

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1