AI Inference Engineer(GPU & KV Cache)

Triune Infomatics Inc

San Jose (CA)

Hybrid

USD 170,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Samsung Cognos seeks an AI Inference Engineer to own the serving stack for LLM inference, focusing on prefill, decode optimization, KV cache offload strategies, and quantization. Work with SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware, and evaluate cross-machine cache sharing protocols.

This is a high-impact role for production-scale AI deployment. The position requires hands-on experience with LLM serving frameworks, GPU programming, and distributed computing concepts.

Qualifications

  • Bachelor’s or Master’s degree in CS/EE or closely related field.
  • Hands-on experience with LLM serving frameworks such as SGLang, vLLM, or TensorRT-LLM.
  • Strong understanding of transformer internals, including attention mechanisms and KV cache.
  • Experience with GPU programming: CUDA, ROCm/HIP, or equivalent.

Responsibilities

  • Design and optimize KV cache offload strategies across a tiered memory hierarchy: GPU HBM, CPU DRAM, and NVMe SSD.
  • Implement and tune PagedAttention, RadixAttention, chunked prefill, and prefix caching within SGLang and vLLM-style serving frameworks.
  • Drive prefill and decode stage optimization to maximize throughput and minimize latency for long-context and multi-user workloads.
  • Apply quantization techniques and speculative decoding to reduce memory footprint and improve inference speed without degrading output quality.
  • Architect and implement tensor parallelism and pipeline parallelism for multi-GPU inference deployments.
  • Tune ROCm kernels for AMD MI300 and MI250 GPUs, adapting CUDA-based patterns to the AMD software stack.
  • Evaluate RDMA and RoCE solutions for cross-node KV cache sharing and assess feasibility in production environments.
  • Collaborate closely with the GPU Software Engineering team on memory-path correctness, data movement latency, and scheduler integration.
  • Profile, benchmark, and iterate on inference stack performance across diverse model architectures and workload profiles.
  • Contribute to design reviews, technical documentation, and team knowledge-sharing sessions.

Skills

GPU Management
LLM Serving
CUDA/ROCm
Python
C++
Distributed Computing
Quantization
Profiling Tools

Education

Bachelor's or Master's in CS/EE

Tools

SGLang
vLLM
TensorRT-LLM
Nsight
ROCProfiler

Job description

AI Inference Engineer ( GPU and KV Cache)

|Duration: 12+ Months |Engagement: Contract

Experience in the following is a MUST

GPU Management

About the Engagement

Samsung Cognos is building a next-generation LLM inference layer in partnership with SGLang, one of the leading open-source serving frameworks in the space. The project addresses one of the hardest problems in large-scale AI deployment: making KV cache memory management fast, efficient, and cost-effective across tiered storage hierarchies at production scale. This is a high-impact engineering engagement where your work will directly influence how AI inference performs for thousands of concurrent users.

Role Overview

As an AI Inference Engineer, you will own the serving stack that powers Samsung's LLM inference layer. Your focus will span prefill and decode optimization, KV cache offload strategies, quantization, speculative decoding, and tensor and pipeline parallelism. You will work within SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware (MI300 and MI250 series), and evaluate RDMA and RoCE protocols for cross-machine cache sharing. This role is algorithm and framework focused: the core question you are answering is whether the model serving logic itself is efficient, scalable, and production-ready.

Work Model: This position follows a Hybrid/Onsite schedule at the Samsung Cognos facility in San Jose, CA. Candidates must be willing and able to work onsite as required by the client.

Key Responsibilities

  • Design and optimize KV cache offload strategies across a tiered memory hierarchy: GPU HBM, CPU DRAM, and NVMe SSD.
  • Implement and tune PagedAttention, RadixAttention, chunked prefill, and prefix caching within SGLang and vLLM-style serving frameworks.
  • Drive prefill and decode stage optimization to maximize throughput and minimize latency for long-context and multi-user workloads.
  • Apply quantization techniques and speculative decoding to reduce memory footprint and improve inference speed without degrading output quality.
  • Architect and implement tensor parallelism and pipeline parallelism for multi-GPU inference deployments.
  • Tune ROCm kernels for AMD MI300 and MI250 GPUs, adapting CUDA-based patterns to the AMD software stack.
  • Evaluate RDMA and RoCE solutions for cross-node KV cache sharing and assess feasibility in production environments.
  • Collaborate closely with the GPU Software Engineering team on memory-path correctness, data movement latency, and scheduler integration.
  • Profile, benchmark, and iterate on inference stack performance across diverse model architectures and workload profiles.
  • Contribute to design reviews, technical documentation, and team knowledge-sharing sessions.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a closely related field.
  • Hands-on experience with LLM serving frameworks such as SGLang, vLLM, or TensorRT-LLM in a production or research setting.
  • Strong understanding of transformer architecture internals, including attention mechanisms and the KV cache.
  • Demonstrated experience with GPU programming: CUDA, ROCm/HIP, or equivalent.
  • Proficiency in Python and C++ for systems-level performance work.
  • Solid grounding in distributed computing concepts including tensor parallelism, pipeline parallelism, and data parallelism.
  • Experience with quantization methods (INT8, FP8, GPTQ, AWQ) and their performance tradeoffs.
  • Ability to profile and optimize inference pipelines using tools such as NSight, ROCProfiler, or equivalent.
  • Strong analytical skills with the ability to translate benchmarking data into actionable engineering decisions.

Preferred Qualifications

  • Direct experience with AMD MI300 or MI250 GPU hardware and the ROCm software ecosystem.
  • Familiarity with speculative decoding frameworks (Medusa, EAGLE, or similar).
  • Exposure to RDMA, RoCE, or InfiniBand for low-latency cross-node communication.
  • Knowledge of NVMe storage characteristics and GPUDirect Storage (GDS) or equivalent technologies.
  • Prior work in a customer-facing or product-integrated inference environment.
  • Open-source contributions to SGLang, vLLM, or related LLM infrastructure projects.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Engineer: Multi-GPU KV Cache & Throughput
LLM Inference Engineer: Multi-GPU KV Cache & Throughput

Triune Infomatics Inc • San Jose (CA)

Hybrid
USD 170,000 - 210,000
Senior Staff Engineer - AI Workloads & Storage
Senior Staff Engineer - AI Workloads & Storage

Engg • San Jose (CA)

On-site
USD 220,000 - 320,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
Senior Research Engineer / Scientist - Storage for LLM
Senior Research Engineer / Scientist - Storage for LLM

ByteDance • Seattle (WA)

On-site
USD 202,160 - 368,220
Medical, dental, vision insurance
401(k) matching
Paid parental leave
+1
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Post-Training Platform Infrastructure Engineer
Post-Training Platform Infrastructure Engineer

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 100,000 - 150,000
Comprehensive benefits
Collaborative work environment
Opportunities for career advancement
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements