Cluster Operations Engineer — MLOps & Orchestration

Comet Cloud

Boston (MA)

Hybrid

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Comet Cloud in the Boston area is hiring a Cluster Operations Engineer to own MLOps and orchestration for GPU clusters, including NVIDIA hardware and software stacks. The role blends platform engineering with customer interaction, focusing on Kubernetes/Slurm orchestration, DCGM, NCCL, and maintaining control plane health while translating workload needs into cluster configurations.

Ideal candidates have production Kubernetes/Slurm experience, model training expertise, and a hands-on approach in

Qualifications

  • Experience running Kubernetes and/or Slurm in production.
  • Expertise in training models, workload orchestration, and optimizing core infrastructure.
  • Familiarity with NVIDIA toolchain or willingness to learn.
  • Ability to monitor and operate control plane infrastructure.
  • Proven ability to write actionable runbooks.

Responsibilities

  • Manage cluster lifecycle and orchestration for GPUs.
  • Configure Kubernetes/Slurm for enterprise workloads.
  • Maintain NVIDIA software stack (DCGM, Base Command Manager, NCCL).
  • Engage with customer ML teams to translate workload needs.
  • Monitor control plane health and performance.

Skills

Kubernetes
Slurm
NVIDIA tools
Model training
Performance tuning
Customer support

Tools

DCGM
Base Command Manager
NCCL
Infiniband

Job description

Cluster Operations Engineer — MLOps & Orchestration (preferably Boston/Dallas; hybrid or remote)

After we deploy a GPU cluster, someone has to make it a great place to actually train models. That's this role.

We build dedicated NVIDIA clusters — GB300/GB200 NVL72, HGX-class B300 nodes, Quantum-3 XDR Infiniband and more for enterprise customers who get a named engineer, not a ticket queue. You'd be that engineer for the platform layer: Kubernetes and Slurm orchestration, the NVIDIA software stack (DCGM, Base Command Manager, NCCL), control plane health, and the direct customer relationship that turns their workload requirements into cluster configuration.

You're a fit if you:

  • - Have run Kubernetes and/or Slurm in production and actually enjoy scheduler configuration
  • - Have a expertise in training models, workload orchestration, optimizing core infrastructure, and resolving performance issues
  • - Know the NVIDIA toolchain, or know one part deeply and want the rest
  • - Can monitor and operate control plane infrastructure without drama
  • - Like sitting with a customer's ML team as much as sitting in a terminal
  • - Write runbooks people actually use

Why here: Lean team, no process layer between you and the machines, and the newest NVIDIA hardware in production. Your customers know your name.

Boston metro or Dallas | Hybrid or remote | Periodic site travel

Comet Compute is an NVIDIA Inception member building to the NVIDIA Cloud Partner reference architecture. Equal opportunity employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cluster Operations Engineer
Cluster Operations Engineer

Comet Cloud • Boston (MA)

Hybrid
USD 140,000 - 200,000
Senior MLOps Engineer — GPU Cluster Orchestration
Senior MLOps Engineer — GPU Cluster Orchestration

Comet Cloud • Boston (MA)

Hybrid
USD 120,000 - 180,000
Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA
Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA

Comet Cloud • Boston (MA)

Hybrid
USD 140,000 - 200,000
Deployment Lead — GPU Cluster Field Deployments
Deployment Lead — GPU Cluster Field Deployments

Comet Cloud • Boston (MA)

On-site
USD 120,000 - 180,000
Hands-on ownership
Access to latest NVIDIA hardware
Deployments with travel
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000