Get more replies from employers
Send a job-specific resume in minutes.
ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer for LLM inference across GPUs and nodes. This role focuses on reducing latency, increasing throughput, and lowering cost for transformer‑based model serving.
You will collaborate with inference and serving teams, optimize eviction policies, memory management, and explore open‑source KV stores or GPU‑aware caching techniques.
ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-performance KV cache layer for LLM inference across GPUs and nodes. This role focuses on reducing latency, increasing throughput, and lowering cost for transformer‑based model serving.
You will collaborate with inference and serving teams, optimize eviction policies, memory management, and explore open‑source KV stores or GPU‑aware caching techniques.