A complete application in a minute — tailored resume and cover letter, ready to send.
Sciforium is seeking a GPU Cluster Engineer to own the software stack for high-performance GPU clusters. You will define production-ready node images, validate hardware with automated tests, and manage fleet upgrades while ensuring consistent, fast, and scalable performance.
You will work with two stakeholder groups—foundation model teams and model serving teams—deploying in a Kubernetes/Slurm environment with advanced NVIDIA/ROCm stacks and IaC tooling.
Sciforium is seeking a GPU Cluster Engineer to own the software stack for high-performance GPU clusters. You will define production-ready node images, validate hardware with automated tests, and manage fleet upgrades while ensuring consistent, fast, and scalable performance.
You will work with two stakeholder groups—foundation model teams and model serving teams—deploying in a Kubernetes/Slurm environment with advanced NVIDIA/ROCm stacks and IaC tooling.