Get more replies from employers
Send a job-specific resume in minutes.
Intel is hiring for an AI inference engineer focusing on optimizing local inference engines (llama.cpp, vLLM) for edge devices and hybrid environments. You will work on KV cache, batching, quantization, and CPU overhead reduction to enable fast, trustworthy local AI solutions.
The role demands deep C++/Python skills, Linux expertise, and experience with LLM inference at scale. This is a hybrid position with multiple US locations including Oregon, Hillsboro, and surrounding areas.
Intel is hiring for an AI inference engineer focusing on optimizing local inference engines (llama.cpp, vLLM) for edge devices and hybrid environments. You will work on KV cache, batching, quantization, and CPU overhead reduction to enable fast, trustworthy local AI solutions.
The role demands deep C++/Python skills, Linux expertise, and experience with LLM inference at scale. This is a hybrid position with multiple US locations including Oregon, Hillsboro, and surrounding areas.