Get more replies from employers
Send a job-specific resume in minutes.
Intelligent Systems Company in London seeks a senior Inference Platform Engineer to own end‑to‑end performance for our heterogeneous inference stack. You will drive KV cache strategies, batching, memory management, and multi‑node scheduling to scale models and accelerate silicon efficiency.
The role requires deep knowledge of LLM internals, strong systems engineering, and hands‑on debugging across GPU and Linux. We offer visa sponsorship and in‑person work at our London office.
Intelligent Systems Company in London seeks a senior Inference Platform Engineer to own end‑to‑end performance for our heterogeneous inference stack. You will drive KV cache strategies, batching, memory management, and multi‑node scheduling to scale models and accelerate silicon efficiency.
The role requires deep knowledge of LLM internals, strong systems engineering, and hands‑on debugging across GPU and Linux. We offer visa sponsorship and in‑person work at our London office.