Get more replies from employers
Send a job-specific resume in minutes.
Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality.
This role focuses on low-latency, privacy-preserving AI workloads run on client devices. You will work across hardware tiers, benchmark performance, and contribute upstream fixes to open-source engines, helping deliver safe, efficient AI at scale while maintaining energy
Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality.
This role focuses on low-latency, privacy-preserving AI workloads run on client devices. You will work across hardware tiers, benchmark performance, and contribute upstream fixes to open-source engines, helping deliver safe, efficient AI at scale while maintaining energy