A complete application in a minute — tailored resume and cover letter, ready to send.
Inferact is seeking a hands-on cluster administration engineer to own and operate high-performance GPU compute infrastructure. You will ensure health, availability, and observability of clusters across neo-cloud and dedicated providers, enabling engineers to build, test, and improve vLLM-powered systems.
You will manage GPU servers, driver health, scheduling, and incident response, partnering with leadership to standardize provisioning and debugging while expanding compute capacity for fast AI
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.
We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.
You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.
Fresh graduates are welcome to apply. Minimum years of work experience: 0.
Monthly salary of S$15,000 to S$30,000, depending on background, skills, and experience, plus equity.
Inferact offers a generous benefits package, including medical, dental, and vision coverage.