LLM Inference Engineer: Multi-GPU KV Cache & Throughput

Triune Infomatics Inc

San Jose (CA)

Hybrid

USD 170,000 - 210,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Samsung Cognos seeks an AI Inference Engineer to own the serving stack for LLM inference, focusing on prefill, decode optimization, KV cache offload strategies, and quantization. Work with SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware, and evaluate cross-machine cache sharing protocols.

This is a high-impact role for production-scale AI deployment. The position requires hands-on experience with LLM serving frameworks, GPU programming, and distributed computing concepts.

Qualifications

  • Bachelor’s or Master’s degree in CS/EE or closely related field.
  • Hands-on experience with LLM serving frameworks such as SGLang, vLLM, or TensorRT-LLM.
  • Strong understanding of transformer internals, including attention mechanisms and KV cache.
  • Experience with GPU programming: CUDA, ROCm/HIP, or equivalent.

Responsibilities

  • Design and optimize KV cache offload strategies across a tiered memory hierarchy: GPU HBM, CPU DRAM, and NVMe SSD.
  • Implement and tune PagedAttention, RadixAttention, chunked prefill, and prefix caching within SGLang and vLLM-style serving frameworks.
  • Drive prefill and decode stage optimization to maximize throughput and minimize latency for long-context and multi-user workloads.
  • Apply quantization techniques and speculative decoding to reduce memory footprint and improve inference speed without degrading output quality.
  • Architect and implement tensor parallelism and pipeline parallelism for multi-GPU inference deployments.
  • Tune ROCm kernels for AMD MI300 and MI250 GPUs, adapting CUDA-based patterns to the AMD software stack.
  • Evaluate RDMA and RoCE solutions for cross-node KV cache sharing and assess feasibility in production environments.
  • Collaborate closely with the GPU Software Engineering team on memory-path correctness, data movement latency, and scheduler integration.
  • Profile, benchmark, and iterate on inference stack performance across diverse model architectures and workload profiles.
  • Contribute to design reviews, technical documentation, and team knowledge-sharing sessions.

Skills

GPU Management
LLM Serving
CUDA/ROCm
Python
C++
Distributed Computing
Quantization
Profiling Tools

Education

Bachelor's or Master's in CS/EE

Tools

SGLang
vLLM
TensorRT-LLM
Nsight
ROCProfiler

Job description

Samsung Cognos seeks an AI Inference Engineer to own the serving stack for LLM inference, focusing on prefill, decode optimization, KV cache offload strategies, and quantization. Work with SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware, and evaluate cross-machine cache sharing protocols.

This is a high-impact role for production-scale AI deployment. The position requires hands-on experience with LLM serving frameworks, GPU programming, and distributed computing concepts.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer(GPU & KV Cache)
AI Inference Engineer(GPU & KV Cache)

Triune Infomatics Inc • San Jose (CA)

Hybrid
USD 170,000 - 210,000
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Principal Software Engineer, Inference
Principal Software Engineer, Inference

Hewlett Packard Enterprise • Fort Collins (CO)

Hybrid
USD 180,000 - 250,000
Health and wellbeing benefits
Professional development programs
Flexible work arrangements
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Sail • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior LLM Storage Systems Engineer & Researcher
Senior LLM Storage Systems Engineer & Researcher

ByteDance • Seattle (WA)

On-site
USD 202,160 - 368,220
Medical, dental, vision insurance
401(k) matching
Paid parental leave
+1
Senior LLM Inference & Serving Engineer
Senior LLM Inference & Serving Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000