A complete application in a minute — tailored resume and cover letter, ready to send.
Cerebras Systems is hiring a Software Engineer to productionize and optimize the GPU serving stack across custom inference APIs, vLLM, and ROCm-based infrastructure. This hands-on role emphasizes performance, observability, and reliability in a disaggregated AI inference environment.
You will work across application, runtime, distributed systems, and hardware layers to boost time-to-first-token, throughput, tail latency, and capacity efficiency on the Cerebras Wafer-Scale Engine.
Cerebras Systems is hiring a Software Engineer to productionize and optimize the GPU serving stack across custom inference APIs, vLLM, and ROCm-based infrastructure. This hands-on role emphasizes performance, observability, and reliability in a disaggregated AI inference environment.
You will work across application, runtime, distributed systems, and hardware layers to boost time-to-first-token, throughput, tail latency, and capacity efficiency on the Cerebras Wafer-Scale Engine.