Senior GPU Cloud K8s Architect

Bitdeer (NASDAQ: BTDR)

San Jose (CA)

On-site

USD 180,000 - 300,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer is building an AI-operated GPU cloud. You will design, deploy, and operate the Kubernetes-based control plane that enables topology-aware scheduling, self-healing, and automated remediation for GPU workloads.

In year one you will optimize multi-tenant clusters, implement BMaaS provisioning, and integrate Prometheus/Grafana monitoring. You will work with Go/Python to develop operators and CRDs that drive autonomous workflows.

Qualifications

  • 5+ years in Kubernetes operations with hands-on GPU workload management
  • Deep expertise with Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation
  • Experience with bare-metal provisioning and lifecycle automation
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Programming skills in Go or Python for operator/CRD development
  • AIOps mindset: automate remediation as part of the control plane
  • Runbook-as-code mindset for executable playbooks

Responsibilities

  • Production Kubernetes clusters optimized for GPU workloads (100-10,000 GPUs)
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies
  • Topology-aware scheduling: GPU locality, NVLink awareness, network affinity
  • CRDs for GPU workload lifecycle management
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security
  • BMaaS provisioning: automated onboarding, lifecycle, reclamation
  • Terraform providers and modules for infrastructure-as-code across GPU clusters
  • SLIs/SLOs for cluster availability and provisioning latency
  • Incident management: runbook automation, escalation, post-incident reviews
  • Monitoring: Prometheus, Grafana, Alertmanager
  • GPU node failure handling: automated drain/cordon/taint and reschedule

Skills

Kubernetes operations
GPU workloads on K8S
Topology-aware scheduling
Multi-tenant platforms
Go or Python

Tools

Terraform
Helm
GitOps
Prometheus
Grafana

Job description

Bitdeer is building an AI-operated GPU cloud. You will design, deploy, and operate the Kubernetes-based control plane that enables topology-aware scheduling, self-healing, and automated remediation for GPU workloads.

In year one you will optimize multi-tenant clusters, implement BMaaS provisioning, and integrate Prometheus/Grafana monitoring. You will work with Go/Python to develop operators and CRDs that drive autonomous workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead GPU Cloud K8s Architect & AIOps Engineer
Lead GPU Cloud K8s Architect & AIOps Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior GPU Cloud Platform Engineer
Senior GPU Cloud Platform Engineer

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 230,000
Senior GPU Fabric Architect for AI Cloud Infra
Senior GPU Fabric Architect for AI Cloud Infra

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Sr GPU Cloud K8S Expert (SRE SME)
Sr GPU Cloud K8S Expert (SRE SME)

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Sr GPU Cloud K8S Expert (SRE SME)
Sr GPU Cloud K8S Expert (SRE SME)

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 230,000
Sr GPU Cloud K8S Expert (SRE SME)
Sr GPU Cloud K8S Expert (SRE SME)

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 300,000
Lead GPU Systems & Fabric Architect for AI Cloud
Lead GPU Systems & Fabric Architect for AI Cloud

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior GPU Systems Architect for High-Performance AI Fabric
Senior GPU Systems Architect for High-Performance AI Fabric

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
K8 Site Reliability SME
K8 Site Reliability SME

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000