Senior Slurm HPC & Kubernetes Scheduler Engineer

Bitdeer Technologies Group

San Jose (CA)

On-site

USD 190,000 - 270,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer Technologies Group seeks a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm across GPU fleets, both bare metal and VM-based nodes, and lead the adoption of the Slinky operator stack for unified GPU resource management.

The role is highly hands-on, customer-facing during onboarding and escalations, with responsibility for multi-tenant scheduling, reliability, and platform-level work across Kubernetes integration and Terraform-driven automation.

Qualifications

  • 8+ years in HPC, systems, or cloud infra engineering.
  • 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Deep hands-on Slurm expertise with slurm.conf, slurmdbd, MUNGE and JWT authentication.
  • Strong GPU and fabric fundamentals: NVIDIA drivers, InfiniBand/RoCE, DCGM.
  • Production Kubernetes experience and Slurm-on-Kubernetes exposure.
  • Go/Python scripting and automation experience.
  • Mature English written and verbal communication.

Responsibilities

  • Design, deploy, and operate production Slurm clusters on bare metal and VM-based GPUs.
  • Lead topology-aware scheduling for GPU fabrics and validate NCCL bandwidth.
  • Own multi-tenant scheduling policy, including partitions, QOS, and reservations.
  • Drive Slinky on Kubernetes and evaluate slurm-bridge for co-scheduling.
  • Manage elastic capacity between Slurm and Kubernetes based on demand.
  • Operate Pyxis/Enroot and container runtimes with correct GPU constraints.
  • Write runbooks, onboard enterprise customers, and mentor engineers.

Skills

HPC engineering
Slurm expertise
Kubernetes experience
Communication in English

Tools

Python
Bash
Go
Terraform
Ansible
Kubernetes
Docker
NVIDIA drivers
DCGM
InfiniBand/RoCE
NCCL
Slurm
Slinky

Job description

Bitdeer Technologies Group seeks a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm across GPU fleets, both bare metal and VM-based nodes, and lead the adoption of the Slinky operator stack for unified GPU resource management.

The role is highly hands-on, customer-facing during onboarding and escalations, with responsibility for multi-tenant scheduling, reliability, and platform-level work across Kubernetes integration and Terraform-driven automation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Slurm & HPC Systems Engineer (GPU/Kubernetes)
Senior Slurm & HPC Systems Engineer (GPU/Kubernetes)

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 250,000
Senior Slurm HPC & Scheduling Engineer
Senior Slurm HPC & Scheduling Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 190,000
Senior Slurm & HPC Cluster Engineer (GPU/AI Infra)
Senior Slurm & HPC Cluster Engineer (GPU/AI Infra)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 150,000 - 190,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 250,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 190,000 - 270,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 190,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 150,000 - 190,000
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity
Senior HPC Scheduler Engineer (LSF/Slurm) - Hybrid & Equity

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Equity
Benefits package
Senior AI Scheduling & Orchestration Engineer
Senior AI Scheduling & Orchestration Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 180,000 - 280,000
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native

Nscale • New York (NY)

On-site
USD 180,000 - 260,000
Bonus
Equity
Medical Insurance
+5