Sr GPU Cloud K8S Expert (SRE SME)

Bitdeer Technologies Group

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer is seeking a seasoned Platform Engineer to run production Kubernetes clusters optimized for GPU workloads at scale. You will design and operate the AIOps-enabled control plane, manage Nvidia GPU tools, and enforce multi-tenant isolation with strong SLIs/SLOs and a robust runbook mindset.

The role requires deep expertise in GPU scheduling, CRD development, and integration of AI frameworks on Kubernetes, with hands-on provisioning of bare-metal resources and GitOps-driven workflows for

Qualifications

  • 5+ years in Kubernetes operations with at least 2 years managing GPU workloads on K8S.
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S.
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for operator/CRD development.
  • AIOps aptitude — automation-first mindset for K8S control plane.

Responsibilities

  • Operate production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Manage Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Implement topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Develop CRDs for GPU workload lifecycle management and integrate AI frameworks on K8S (Slurm, Ray, Kubeflow).
  • Maintain multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Handle BMaaS provisioning and tenant onboarding lifecycle; automate provisioning/remediation workflows.
  • Maintain Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • Define and monitor SLIs/SLOs for cluster availability and job completion latency.
  • Oversee incident management: runbook automation, escalation, post-incident reviews.
  • Monitor stack: Prometheus, Grafana, Alertmanager; optimize for reliability at scale.

Skills

Kubernetes operations
GPU workloads on K8S
NVIDIA GPU operator
Topology-aware scheduling
Multi-tenant K8S
Bare-metal provisioning
Terraform
GitOps (ArgoCD/Flux)
Prometheus & Grafana
Go or Python

Tools

Ironic
MAAS
ArgoCD
Flux

Job description

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/

Position Overview

You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own
  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
Feed the AIOps substrate
  • The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.
What success looks like in year 1
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.
Job Requirement:
  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

K8 Site Reliability SME
K8 Site Reliability SME

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Sr. GPU Cloud Storage Solutions Expert (SRE SME)
Sr. GPU Cloud Storage Solutions Expert (SRE SME)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 170,000 - 250,000
SRE Platform Software Engineer (Early Career / Temporary)
SRE Platform Software Engineer (Early Career / Temporary)

Bitdeer Technologies Group • Austin (TX)

On-site
USD 110,000 - 140,000
SRE Monitoring Platform Software Engineer (Entry Level)
SRE Monitoring Platform Software Engineer (Entry Level)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 90,000 - 125,000