Senior Slurm & HPC Systems Engineer (GPU/Kubernetes)

Bitdeer (NASDAQ: BTDR)

San Jose (CA)

On-site

USD 180,000 - 250,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Bitdeer is seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across a GPU fleet. You will be the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability, driving adoption of the Slinky operator stack to share GPU pools with Kubernetes workloads.

The role is hands-on, customer-facing during onboarding and escalations, and sets the engineering standards for the platform team.

Qualifications

  • 8+ years in HPC or cloud infra engineering.
  • Deep hands-on Slurm expertise: slurm.conf, slurmdbd, slurmrestd, MUNGE/SACK, JWT authentication.

Responsibilities

  • Design, deploy, and operate production Slurm clusters on bare metal and VMs.
  • Lead Slinky on Kubernetes implementation including NodeSet and Accounting resources.
  • Enable elastic capacity between Slurm and Kubernetes for dynamic workloads.
  • Maintain cluster health with monitoring, diagnostics, and reliability engineering.
  • Provide customer onboarding, runbooks, and tenant-facing docs.

Skills

Slurm expertise
HPC / Cloud
Kubernetes
Python
Go
Terraform/Ansible

Education

Bachelor's degree in Computer Science/Engineering

Tools

slurm-operator
slurm-bridge
Kubernetes
Terraform
Ansible

Job description

Bitdeer is seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across a GPU fleet. You will be the single technical owner of Slurm cluster architecture, multi-tenant scheduling policy, and cluster reliability, driving adoption of the Slinky operator stack to share GPU pools with Kubernetes workloads.

The role is hands-on, customer-facing during onboarding and escalations, and sets the engineering standards for the platform team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Slurm HPC & Kubernetes Scheduler Engineer
Senior Slurm HPC & Kubernetes Scheduler Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 190,000 - 270,000
Senior Slurm HPC & Scheduling Engineer
Senior Slurm HPC & Scheduling Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 190,000
Senior Slurm & HPC Cluster Engineer (GPU/AI Infra)
Senior Slurm & HPC Cluster Engineer (GPU/AI Infra)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 150,000 - 190,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 190,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 150,000 - 190,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 250,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 190,000 - 270,000
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native
Senior HPC Systems Engineer — Slurm, GPU, Cloud-Native

Nscale • New York (NY)

On-site
USD 180,000 - 260,000
Bonus
Equity
Medical Insurance
+5
Senior Backend Engineer: Kubernetes & Slurm Platform
Senior Backend Engineer: Kubernetes & Slurm Platform

Lightning-Ai • New York (NY)

Hybrid
USD 180,000 - 250,000
Health coverage
Equity / RSUs
401(k) matching
+8
Senior Backend Engineer — Kubernetes & Slurm Platform
Senior Backend Engineer — Kubernetes & Slurm Platform

Lightning AI • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Comprehensive Health Coverage
Meaningful Equity
401(k) matching
+8