Senior Edge Inference Optimization Engineer

Intel Corporation

Santa Clara (CA)

Hybrid

USD 195,000 - 361,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs.

You’ll benchmark across hardware tiers and contribute patches to open‑source engines, aligning with Intel’s mission to improve AI safety, privacy, and efficiency on local devices.

Qualifications

  • BS/MS in CS, EE, Math or related STEM field.
  • 8+ years software development background.
  • Strong in C++ and/or Python; comfortable reading systems‑level code.
  • Experience with LLM inference (attention, KV cache, decoding).
  • Experience profiling and optimizing real performance problems (CPU or GPU).
  • Linux, build systems, and low‑level debugging expertise.
  • Hands‑on with llama.cpp, vLLM, ggml, or similar engines.
  • Experience with GPU / accelerator programming (Vulkan, CUDA, SYCL, Metal).
  • Familiarity with quantization formats and their quality trade‑offs.
  • Open‑source contributions to inference engines.

Responsibilities

  • Profile and optimize local inference for latency, throughput, and memory on edge hardware.
  • Tune KV cache, continuous batching, and scheduling for interactive workloads.
  • Drive quantization strategy and validate quality impact with the Post‑Training team.
  • Cut CPU overhead and improve engine startup, model load, and lifecycle (start / stop / health).
  • Benchmark across hardware tiers and publish honest performance comparisons.
  • Upstream fixes and patches to open‑source engines where it helps us.

Skills

C++
Python
LLM inference
Profiling & optimization
Linux
Low-level debugging
Systems programming
Vulkan/CUDA/SYCL/Metal
Quantization formats
Open-source contributions

Education

BS/MS in CS, EE, Math or related STEM field

Tools

llama.cpp
vLLM
ggml
Vulkan
CUDA
SYCL
Metal

Job description

Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs.

You’ll benchmark across hardware tiers and contribute patches to open‑source engines, aligning with Intel’s mission to improve AI safety, privacy, and efficiency on local devices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Edge AI Inference Optimization Engineer
Senior Edge AI Inference Optimization Engineer

PVH (Tommy Hilfiger/Calvin Klein) • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Intel Benefits
Senior Edge Inference Optimization Engineer
Senior Edge Inference Optimization Engineer

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Intel Benefits
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel Corporation • Santa Clara (CA)

Hybrid
USD 195,000 - 361,000
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Folsom (CA)

Hybrid
USD 195,000 - 362,000
Edge-Optimized AI Inference Kernel Engineer
Edge-Optimized AI Inference Kernel Engineer

Framework Ventures • United States

Remote
USD 180,000 - 260,000
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior GPU AI Platform Engineer — Edge Inference (Equity)
Senior GPU AI Platform Engineer — Edge Inference (Equity)

NVIDIA AI • Seattle (WA)

On-site
USD 224,000 - 431,250
Equity
Benefits