Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR)

Austin (TX)

On-site

USD 140,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Bitdeer Technologies Group is building an AI-operated GPU cloud where reliability is the product. You will own end-to-end reliability from onboarding to workload execution and incident response, ensuring tenants experience high availability and predictable performance.

As the SRE lead for the GPU cloud, you will design observability, automation, and operational practices for a 10,000-GPU environment, enabling self-healing and scalable workflows that keep tenant workloads on track.

Qualifications

  • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Customer-facing cloud service experience - defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation / operator development.
  • AIOps aptitude - you view the control plane as an execution surface for automated remediation, not just a scheduler.
  • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

Responsibilities

  • Ensure end-to-end reliability of the GPU cloud service for tenants.
  • Manage production Kubernetes clusters for GPU workloads at scale.
  • Implement topology-aware scheduling and multi-tenant isolation policies.
  • Oversee BMaaS provisioning and tenant lifecycle management.
  • Define and publish customer-facing SLAs/SLOs and drive error-budget prioritization.
  • Build self-service observability for customers to check status and quotas.
  • Develop runbooks and automation to enable automated remediation.

Skills

SRE Operations
Kubernetes
GPU Scheduling
Terraform
Go or Python
GitOps

Tools

NVIDIA GPU operator
Device plugin
MIG
ArgoCD
Flux
Prometheus
Grafana
Ironic/MAAS
Terraform
Helm

Job description

Bitdeer Technologies Group is building an AI-operated GPU cloud where reliability is the product. You will own end-to-end reliability from onboarding to workload execution and incident response, ensuring tenants experience high availability and predictable performance.

As the SRE lead for the GPU cloud, you will design observability, automation, and operational practices for a 10,000-GPU environment, enabling self-healing and scalable workflows that keep tenant workloads on track.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Senior SRE - Customer-Facing GPU Cloud Reliability
Senior SRE - Customer-Facing GPU Cloud Reliability

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000
Junior SRE: Monitoring Platform for GPU Cloud
Junior SRE: Monitoring Platform for GPU Cloud

Bitdeer Technologies Group • Austin (TX)

On-site
USD 70,000 - 100,000
Junior SRE: GPU Cloud Observability Platform
Junior SRE: GPU Cloud Observability Platform

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 90,000 - 125,000
Entry-Level SRE: Observability Platform for GPU Cloud
Entry-Level SRE: Observability Platform for GPU Cloud

Bitdeer • San Jose (CA)

On-site
USD 110,000 - 150,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000
Remote Senior SRE: GPU Cloud, Automation & Reliability
Remote Senior SRE: GPU Cloud, Automation & Reliability

asobbi • United States

On-site
USD 150,000 - 190,000
Equity
Bonus
Benefits
Senior DevOps Engineer — CI/CD for AI GPU Cloud
Senior DevOps Engineer — CI/CD for AI GPU Cloud

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 240,000