Remote Cluster Ops Engineer - Kubernetes, Slurm, NVIDIA

Comet Cloud

Boston (MA)

Hybrid

USD 140,000 - 200,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Comet Cloud in Boston is seeking a hands-on Platform Engineer to manage production Kubernetes and Slurm clusters for NVIDIA workloads. You will own scheduling policy, run the NVIDIA toolchain (DCGM, Base Command Manager, NCCL), and ensure the control plane is reliable while collaborating with ML teams you onboard.

You’ll work directly with customers, turning training requirements into a configured cluster, with a focus on performance, monitoring, and recoverability.

Responsibilities

  • Kubernetes and Slurm orchestration in production
  • Operate NVIDIA toolchain (DCGM, Base Command Manager, NCCL) on the latest hardware
  • Onboard ML teams and tailor platform to their workloads
  • Maintain a boring, monitored, recoverable control plane

Job description

Comet Cloud in Boston is seeking a hands-on Platform Engineer to manage production Kubernetes and Slurm clusters for NVIDIA workloads. You will own scheduling policy, run the NVIDIA toolchain (DCGM, Base Command Manager, NCCL), and ensure the control plane is reliable while collaborating with ML teams you onboard.

You’ll work directly with customers, turning training requirements into a configured cluster, with a focus on performance, monitoring, and recoverability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cluster Operations Engineer — MLOps & Orchestration
Cluster Operations Engineer — MLOps & Orchestration

Comet Cloud • Boston (MA)

Hybrid
USD 120,000 - 180,000
Senior MLOps Engineer — GPU Cluster Orchestration
Senior MLOps Engineer — GPU Cluster Orchestration

Comet Cloud • Boston (MA)

Hybrid
USD 120,000 - 180,000
Cluster Operations Engineer
Cluster Operations Engineer

Comet Cloud • Boston (MA)

Hybrid
USD 140,000 - 200,000
Senior Cloud-Native AI Data Center Engineer Kubernetes/Slurm
Senior Cloud-Native AI Data Center Engineer Kubernetes/Slurm

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Engineer
Senior Kubernetes Engineer

NMC2 • Dallas (TX)

On-site
USD 120,000 - 160,000
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)
Senior AI Datacenter Software Engineer (Kubernetes/Slurm)

BranchFactor • Austin (TX), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Engineer – GPU & HPC Platform
Senior Kubernetes Engineer – GPU & HPC Platform

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 130,000 - 180,000
Company-Paid lunch stipend
Medical, dental, vision benefits
401(k) matching up to 6%
+1
Senior Kubernetes Platform Engineer - GitOps & Cloud
Senior Kubernetes Platform Engineer - GitOps & Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Comprehensive benefits
On-Site HPC/AI Cluster Deployment Lead
On-Site HPC/AI Cluster Deployment Lead

Comet Cloud • Boston (MA)

Hybrid
USD 100,000 - 150,000
On-Site Field Engineer: GPU Clusters & Rack Hardware
On-Site Field Engineer: GPU Clusters & Rack Hardware

Comet Cloud • Boston (MA)

Hybrid
USD 90,000 - 130,000
Hands-on hardware exposure
Growth and NVIDIA training
Lean team environment