Senior Edge AI Inference Optimization Engineer

PVH (Tommy Hilfiger/Calvin Klein)

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Intel Benefits

Job summary

Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality.

This role focuses on low-latency, privacy-preserving AI workloads run on client devices. You will work across hardware tiers, benchmark performance, and contribute upstream fixes to open-source engines, helping deliver safe, efficient AI at scale while maintaining energy

Qualifications

  • BS/MS in CS/EE/Math or related STEM field.
  • 8+ years software development background.
  • Strong in C++ and/or Python; reading systems-level code.
  • Experience with LLM inference and KV cache.

Responsibilities

  • Profile and optimize local inference for latency and throughput on edge hardware.
  • Tune KV cache, batching, and scheduling for interactive workloads.
  • Drive quantization strategy and validate quality impact.
  • Cut CPU overhead and improve engine startup, model load, and lifecycle.
  • Benchmark across hardware tiers and publish honest performance comparisons.
  • Upstream fixes and patches to open-source engines where it helps us

Skills

C++
Python
LLM inference
Linux
Profiling
Optimization

Education

BS/MS in CS/EE/Math or related STEM field

Tools

llama.cpp
vLLM
ggml
CUDA

Job description

Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality.

This role focuses on low-latency, privacy-preserving AI workloads run on client devices. You will work across hardware tiers, benchmark performance, and contribute upstream fixes to open-source engines, helping deliver safe, efficient AI at scale while maintaining energy

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Edge Inference Optimization Engineer
Senior Edge Inference Optimization Engineer

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Senior Edge Inference Optimization Engineer
Senior Edge Inference Optimization Engineer

Intel Corporation • Santa Clara (CA)

Hybrid
USD 195,000 - 361,000
Senior Edge AI Inference Engineer
Senior Edge AI Inference Engineer

Intel • Santa Clara (CA)

Hybrid
USD 195,000 - 362,000
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

PVH (Tommy Hilfiger/Calvin Klein) • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Intel Benefits
High-Performance AI Inference Engineer
High-Performance AI Inference Engineer

Relha LLC • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Edge-Optimized AI Inference Kernel Engineer
Edge-Optimized AI Inference Kernel Engineer

Framework Ventures • United States

Remote
USD 180,000 - 260,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Santa Clara (CA)

Hybrid
USD 195,000 - 362,000
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel • Phoenix (AZ)

Hybrid
USD 195,000 - 362,000
Stock bonuses
Health insurance
Retirement plan
+1
Sr. Inference Optimization Engineer (local / edge runtime)
Sr. Inference Optimization Engineer (local / edge runtime)

Intel Corporation • Santa Clara (CA)

Hybrid
USD 195,000 - 361,000