Turn this role into an interview — a resume and cover letter built around what this employer wants.
Samsung Cognos seeks an AI Inference Engineer to own the serving stack for LLM inference, focusing on prefill, decode optimization, KV cache offload strategies, and quantization. Work with SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware, and evaluate cross-machine cache sharing protocols.
This is a high-impact role for production-scale AI deployment. The position requires hands-on experience with LLM serving frameworks, GPU programming, and distributed computing concepts.
AI Inference Engineer ( GPU and KV Cache)
|Duration: 12+ Months |Engagement: Contract
Experience in the following is a MUST
GPU Management
About the Engagement
Samsung Cognos is building a next-generation LLM inference layer in partnership with SGLang, one of the leading open-source serving frameworks in the space. The project addresses one of the hardest problems in large-scale AI deployment: making KV cache memory management fast, efficient, and cost-effective across tiered storage hierarchies at production scale. This is a high-impact engineering engagement where your work will directly influence how AI inference performs for thousands of concurrent users.
Role Overview
As an AI Inference Engineer, you will own the serving stack that powers Samsung's LLM inference layer. Your focus will span prefill and decode optimization, KV cache offload strategies, quantization, speculative decoding, and tensor and pipeline parallelism. You will work within SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware (MI300 and MI250 series), and evaluate RDMA and RoCE protocols for cross-machine cache sharing. This role is algorithm and framework focused: the core question you are answering is whether the model serving logic itself is efficient, scalable, and production-ready.
Work Model: This position follows a Hybrid/Onsite schedule at the Samsung Cognos facility in San Jose, CA. Candidates must be willing and able to work onsite as required by the client.
Key Responsibilities
Required Qualifications
Preferred Qualifications