Senior Edge Inference Optimization Engineer

Intel

Phoenix (AZ)

Hybrid

USD 195,200 - 361,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock bonuses
Health insurance
Retirement plan
Paid vacation

Job summary

Intel is seeking a senior software engineer to optimize AI inference for hybrid US environments. You will advance local and edge inference engines (llama.cpp, vLLM) on PC/edge hardware, focusing on latency, memory, and energy efficiency.

Collaborating with teams, you will push quantization, KV cache strategies, and CPU/GPU tuning to deliver strong performance while upholding Intel's commitment to trusted AI and privacy by design.

Qualifications

  • BS/MS in CS, EE, Math or related STEM field.
  • 8+ years software development background.
  • Strong in C++ and/or Python; comfortable reading systems-level code.
  • Experience with LLM inference. (attention, KV cache, decoding)
  • Experience profiling and optimizing real performance problems (CPU or GPU) and can prove the speedup
  • Linux, build systems, and low-level debugging expertise.

Responsibilities

  • Profile and optimize local inference (llama.cpp-vulkan and vLLM) for latency, throughput, and memory on edge hardware
  • Tune KV cache, continuous batching, and scheduling for interactive agent workloads
  • Drive quantization strategy (GGUF / AWQ / GPTQ) and validate quality impact with the Post-Training team
  • Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health)
  • Benchmark across hardware tiers and publish honest performance comparisons
  • Upstream fixes and patches to open-source engines where it helps us

Skills

C++
Python
LLM inference
Profiling
Linux
Low-level debugging

Education

Bachelor's or Master's in CS/EE/Math or related STEM

Tools

llama.cpp
vLLM
ggml
Vulkan
CUDA
SYCL
oneAPI
Metal

Job description

Intel is seeking a senior software engineer to optimize AI inference for hybrid US environments. You will advance local and edge inference engines (llama.cpp, vLLM) on PC/edge hardware, focusing on latency, memory, and energy efficiency.

Collaborating with teams, you will push quantization, KV cache strategies, and CPU/GPU tuning to deliver strong performance while upholding Intel's commitment to trusted AI and privacy by design.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Edge AI Inference Optimization Engineer
Senior Edge AI Inference Optimization Engineer

PVH (Tommy Hilfiger/Calvin Klein) • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Intel Benefits
Senior Edge Inference Optimization Engineer
Senior Edge Inference Optimization Engineer

Intel Corporation • Santa Clara (CA)

Hybrid
USD 195,000 - 361,000
Senior Edge AI Inference Engineer
Senior Edge AI Inference Engineer

Intel • Santa Clara (CA)

Hybrid
USD 195,000 - 362,000
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

PVH (Tommy Hilfiger/Calvin Klein) • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Intel Benefits
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Santa Clara (CA)

Hybrid
USD 195,000 - 362,000
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Folsom (CA)

Hybrid
USD 195,000 - 362,000
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel Corporation • Santa Clara (CA)

Hybrid
USD 195,000 - 361,000
Senior GPU AI Platform Engineer — Edge Inference (Equity)
Senior GPU AI Platform Engineer — Edge Inference (Equity)

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits