An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Comet Cloud in the Boston area is hiring a Cluster Operations Engineer to own MLOps and orchestration for GPU clusters, including NVIDIA hardware and software stacks. The role blends platform engineering with customer interaction, focusing on Kubernetes/Slurm orchestration, DCGM, NCCL, and maintaining control plane health while translating workload needs into cluster configurations.
Ideal candidates have production Kubernetes/Slurm experience, model training expertise, and a hands-on approach in
Cluster Operations Engineer — MLOps & Orchestration (preferably Boston/Dallas; hybrid or remote)
After we deploy a GPU cluster, someone has to make it a great place to actually train models. That's this role.
We build dedicated NVIDIA clusters — GB300/GB200 NVL72, HGX-class B300 nodes, Quantum-3 XDR Infiniband and more for enterprise customers who get a named engineer, not a ticket queue. You'd be that engineer for the platform layer: Kubernetes and Slurm orchestration, the NVIDIA software stack (DCGM, Base Command Manager, NCCL), control plane health, and the direct customer relationship that turns their workload requirements into cluster configuration.
You're a fit if you:
Why here: Lean team, no process layer between you and the machines, and the newest NVIDIA hardware in production. Your customers know your name.
Boston metro or Dallas | Hybrid or remote | Periodic site travel
Comet Compute is an NVIDIA Inception member building to the NVIDIA Cloud Partner reference architecture. Equal opportunity employer.