Cluster Operations Engineer

Comet Cloud

Boston (MA)

Hybrid

USD 140,000 - 200,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Comet Cloud in Boston is seeking a hands-on Platform Engineer to manage production Kubernetes and Slurm clusters for NVIDIA workloads. You will own scheduling policy, run the NVIDIA toolchain (DCGM, Base Command Manager, NCCL), and ensure the control plane is reliable while collaborating with ML teams you onboard.

You’ll work directly with customers, turning training requirements into a configured cluster, with a focus on performance, monitoring, and recoverability.

Responsibilities

  • Kubernetes and Slurm orchestration in production
  • Operate NVIDIA toolchain (DCGM, Base Command Manager, NCCL) on the latest hardware
  • Onboard ML teams and tailor platform to their workloads
  • Maintain a boring, monitored, recoverable control plane

Job description

Boston Metro, MA · Hybrid or Remote · Full-time

Our customers get dedicated NVIDIA clusters and a named engineer who knows their workload, not a ticket queue. This is that engineer, for everything above the metal: Kubernetes and Slurm orchestration, the NVIDIA software stack, control plane health, and the customer relationship that turns "here's what we're training" into a cluster configured to do it well.

This role is for you if
  • You've operated Kubernetes or Slurm in production and have opinions about scheduler policy.
  • You want to run the NVIDIA toolchain (DCGM, Base Command Manager, NCCL) on the newest hardware NVIDIA makes.
  • You like the customer side: onboarding an ML team, learning what they're building, and shaping the platform around it.
  • You believe a control plane should be boring, monitored, and recoverable.
What we promise you
  • Direct ownership, with no layer between you and the clusters or the customers on them.
  • A fleet spanning GB300 NVL72 to H100, on Quantum-3 XDR InfiniBand, scaling toward 9,552 GPUs.
  • Customers whose training runs you'll know by name.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cluster Operations Engineer — MLOps & Orchestration
Cluster Operations Engineer — MLOps & Orchestration

Comet Cloud • Boston (MA)

Hybrid
USD 120,000 - 180,000
Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA
Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA

Comet Cloud • Boston (MA)

Hybrid
USD 140,000 - 200,000
Deployment Lead
Deployment Lead

Comet Cloud • Boston (MA)

Hybrid
USD 100,000 - 150,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Deployment Lead
Deployment Lead

Comet Cloud • Dallas (TX), Northern (KY)

Hybrid
USD 100,000 - 160,000
Deployment Lead — GPU Cluster Field Deployments
Deployment Lead — GPU Cluster Field Deployments

Comet Cloud • Boston (MA)

On-site
USD 120,000 - 180,000
Hands-on ownership
Access to latest NVIDIA hardware
Deployments with travel
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Field Engineer
Field Engineer

Comet Cloud • Boston (MA)

On-site
USD 90,000 - 130,000
Hands-on hardware exposure
Growth and NVIDIA training
Lean team environment
Senior MLOps Engineer — GPU Cluster Orchestration
Senior MLOps Engineer — GPU Cluster Orchestration

Comet Cloud • Boston (MA)

Hybrid
USD 120,000 - 180,000
Field Engineer
Field Engineer

Comet Cloud • Dallas (TX), Northern (KY)

Hybrid
USD 70,000 - 110,000