Senior Edge Inference Optimization Engineer

Intel Corporation

Santa Clara, Northern (CA, KY)

Hybrid

USD 217,000 - 361,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intel is seeking an Sr. Inference Optimization Engineer to accelerate models on edge devices. You will optimize inference engines for constrained hardware, including profiling, KV cache tuning, batching, and quantization.

This role focuses on fast, efficient local AI with privacy and low latency. You will work across iGPU/CPU paths, improve engine startup, and contribute fixes to open-source engines. A strong background in C++, Python, and LLM inference is required, with hybrid on-site work in

Qualifications

  • BS/MS in CS, EE, Math or related STEM field.
  • 8+ years software development background.
  • Strong in C++ and/or Python; comfortable reading systems-level code.
  • Experience with LLM inference (attention, KV cache, decoding).
  • Experience profiling and optimizing real performance problems (CPU or GPU).
  • Linux, build systems, and low-level debugging expertise.

Responsibilities

  • Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware.
  • Tune KV cache, continuous batching, and scheduling for interactive workloads.
  • Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team.
  • Cut CPU overhead and improve engine startup, model load, and lifecycle management.
  • Benchmark across hardware tiers and publish performance comparisons.
  • Upstream fixes and patches to open-source engines where it helps us.

Skills

C++
Python
LLM inference
Profiling
Performance optimization

Education

BS/MS in CS, EE, Math

Tools

llama.cpp
vLLM
ggml

Job description

Intel is seeking an Sr. Inference Optimization Engineer to accelerate models on edge devices. You will optimize inference engines for constrained hardware, including profiling, KV cache tuning, batching, and quantization.

This role focuses on fast, efficient local AI with privacy and low latency. You will work across iGPU/CPU paths, improve engine startup, and contribute fixes to open-source engines. A strong background in C++, Python, and LLM inference is required, with hybrid on-site work in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Edge Inference Optimization Engineer
Senior Edge Inference Optimization Engineer

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Senior Edge AI Inference Engineer
Senior Edge AI Inference Engineer

Intel • Hillsboro (OR)

Hybrid
USD 195,200 - 361,200
Hybrid work model
Competitive compensation
Senior Edge AI/ML Inference Engineer
Senior Edge AI/ML Inference Engineer

Socket.dev • Austin (TX)

On-site
USD 150,000 - 210,000
Competitive salary
ACS Equity Package
Health, Dental, Vision Insurance
+1
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Hillsboro (OR)

Hybrid
USD 195,200 - 361,200
Hybrid work model
Competitive compensation
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Folsom (CA)

Hybrid
USD 195,000 - 362,000
Edge AI Optimization Engineer — LLMs & Inference
Edge AI Optimization Engineer — LLMs & Inference

Artha Nexgen • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 190,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior Edge AI Systems Engineer - Remote
Senior Edge AI Systems Engineer - Remote

Visa Hunt • United States

On-site
USD 100,000 - 150,000