Senior SRE & Automation Architect — GPU Cloud, Remote

Bitdeer Technologies Group

San Jose (CA)

Remote

USD 180,000 - 260,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Bitdeer Technologies Group in the United States is hiring a Senior SRE to own the reliability of the customer-facing GPU cloud service end-to-end, from onboarding to post-incident recovery. You will design observability, automation, and operational practices for a 10,000-GPU cloud, ensuring high availability and dependable tenant experience.

You will operate production Kubernetes clusters, NVIDIA GPU operator, device plugin, MIG, and multi-tenant isolation; build self-service observability,

Qualifications

  • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management (NVIDIA GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Customer-facing cloud service experience — defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation / operator development.
  • AIOps aptitude — you view the control plane as an execution surface for automated remediation, not just a scheduler.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

Responsibilities

  • End-to-end reliability of the customer-facing GPU cloud service — availability, job completion, provisioning latency, and tenant experience.
  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs) as the runtime substrate for customer workloads.
  • Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity — placing customer jobs on the right hardware.
  • Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
  • BMaaS: automated provisioning, tenant handoff, lifecycle, and reclamation.
  • SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
  • Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
  • Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty — tenant-aware dashboards and alerting.
  • GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling — minimizing customer-visible impact.
  • Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
  • Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.

Skills

SRE / cloud operations
Kubernetes operations
GPU workloads at scale
Multi-tenant cloud platforms
Incident management
Terraform
Helm
GitOps (ArgoCD/Flux)
Prometheus & Grafana
Go or Python
Runbook automation

Tools

NVIDIA GPU operator
Ironic/MAAS
ArgoCD/Flux
Prometheus
Grafana

Job description

Bitdeer Technologies Group in the United States is hiring a Senior SRE to own the reliability of the customer-facing GPU cloud service end-to-end, from onboarding to post-incident recovery. You will design observability, automation, and operational practices for a 10,000-GPU cloud, ensuring high availability and dependable tenant experience.

You will operate production Kubernetes clusters, NVIDIA GPU operator, device plugin, MIG, and multi-tenant isolation; build self-service observability,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Senior SRE - GPU Cloud Reliability & Automation
Senior SRE - GPU Cloud Reliability & Automation

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Senior SRE - Customer-Facing GPU Cloud Reliability
Senior SRE - Customer-Facing GPU Cloud Reliability

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000
Senior SRE: Customer-Facing GPU Cloud Reliability
Senior SRE: Customer-Facing GPU Cloud Reliability

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior GPU Cloud Platform Engineer
Senior GPU Cloud Platform Engineer

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 230,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 180,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer Technologies Group • San Jose (CA)

Remote
USD 180,000 - 260,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer • San Jose (CA)

On-site
USD 180,000 - 260,000
Sr SRE & Automation Engineer (Customer Facing)
Sr SRE & Automation Engineer (Customer Facing)

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000