A complete application in a minute — tailored resume and cover letter, ready to send.
Comet Cloud in Boston is seeking a hands-on Platform Engineer to manage production Kubernetes and Slurm clusters for NVIDIA workloads. You will own scheduling policy, run the NVIDIA toolchain (DCGM, Base Command Manager, NCCL), and ensure the control plane is reliable while collaborating with ML teams you onboard.
You’ll work directly with customers, turning training requirements into a configured cluster, with a focus on performance, monitoring, and recoverability.
Boston Metro, MA · Hybrid or Remote · Full-time
Our customers get dedicated NVIDIA clusters and a named engineer who knows their workload, not a ticket queue. This is that engineer, for everything above the metal: Kubernetes and Slurm orchestration, the NVIDIA software stack, control plane health, and the customer relationship that turns "here's what we're training" into a cluster configured to do it well.