Get more replies from employers
Send a job-specific resume in minutes.
Bitdeer is seeking a Senior Inference Runtime Engineer to own the performance-critical serving layer of the MaaS platform. This role focuses on making self-hosted LLMs faster, cheaper, and more stable by optimizing the runtime stack behind OpenAI- and Anthropic-compatible APIs.
You will optimize runtimes, profile bottlenecks, and lead model onboarding while collaborating with SRE and performance teams to deliver reliable, low-latency inference services at scale.
Bitdeer is seeking a Senior Inference Runtime Engineer to own the performance-critical serving layer of the MaaS platform. This role focuses on making self-hosted LLMs faster, cheaper, and more stable by optimizing the runtime stack behind OpenAI- and Anthropic-compatible APIs.
You will optimize runtimes, profile bottlenecks, and lead model onboarding while collaborating with SRE and performance teams to deliver reliable, low-latency inference services at scale.