An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Samsung Cognos seeks an AI Inference Engineer to own the serving stack for LLM inference, focusing on prefill, decode optimization, KV cache offload strategies, and quantization. Work with SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware, and evaluate cross-machine cache sharing protocols.
This is a high-impact role for production-scale AI deployment. The position requires hands-on experience with LLM serving frameworks, GPU programming, and distributed computing concepts.
Samsung Cognos seeks an AI Inference Engineer to own the serving stack for LLM inference, focusing on prefill, decode optimization, KV cache offload strategies, and quantization. Work with SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware, and evaluate cross-machine cache sharing protocols.
This is a high-impact role for production-scale AI deployment. The position requires hands-on experience with LLM serving frameworks, GPU programming, and distributed computing concepts.