Get more replies from employers
Send a job-specific resume in minutes.
Intel is seeking an Sr. Inference Optimization Engineer to accelerate models on edge devices. You will optimize inference engines for constrained hardware, including profiling, KV cache tuning, batching, and quantization.
This role focuses on fast, efficient local AI with privacy and low latency. You will work across iGPU/CPU paths, improve engine startup, and contribute fixes to open-source engines. A strong background in C++, Python, and LLM inference is required, with hybrid on-site work in
Intel is seeking an Sr. Inference Optimization Engineer to accelerate models on edge devices. You will optimize inference engines for constrained hardware, including profiling, KV cache tuning, batching, and quantization.
This role focuses on fast, efficient local AI with privacy and low latency. You will work across iGPU/CPU paths, improve engine startup, and contribute fixes to open-source engines. A strong background in C++, Python, and LLM inference is required, with hybrid on-site work in