Get more replies from employers
Send a job-specific resume in minutes.
Callosum is seeking a seasoned ML inference engineer to push heterogeneous hardware orchestration beyond GPUs. You will work on SGLang and vLLM, extending them to run efficiently across diverse accelerators, focusing on scheduling, memory and execution.
Based in London, you will collaborate with an Accelerator Systems Software engineer, contribute upstream, and maintain internal forks while scaling production inference on mixed hardware.
Callosum is seeking a seasoned ML inference engineer to push heterogeneous hardware orchestration beyond GPUs. You will work on SGLang and vLLM, extending them to run efficiently across diverse accelerators, focusing on scheduling, memory and execution.
Based in London, you will collaborate with an Accelerator Systems Software engineer, contribute upstream, and maintain internal forks while scaling production inference on mixed hardware.