Sr. SRE & Automation Engineer (Customer facing)

Bitdeer Technologies Group

Singapore

On-site

SGD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Welfare benefits
Training and mentoring

Job summary

Bitdeer Technologies Group is seeking an experienced Site Reliability/Cloud Operations Engineer to own the reliability of our GPU cloud service. You will run production Kubernetes clusters for GPU workloads, implement multi-tenant policies, and drive incident response with strong automation and observability.

The role emphasizes end-to-end SLAs, proactive remediation, and collaboration with customer success to improve platform reliability for external tenants.

Qualifications

  • 5+ years in SRE / cloud operations with GPU workloads
  • Deep understanding of Kubernetes operations and GPU workload management
  • Experience building multi-tenant cloud platforms with strong isolation guarantees
  • Ability to define and operate around customer SLAs/SLOs and incident communications
  • Proficiency in Go or Python for automation/operator development
  • Experience with bare-metal provisioning and lifecycle automation (Ironic/MAAS or similar)

Responsibilities

  • Ensure end-to-end reliability of the GPU cloud service.
  • Manage production Kubernetes clusters for GPU workloads (100–10,000 GPUs).
  • Implement Nvidia GPU operator, device plugin, MIG, time-slicing, and multi-tenant policies.
  • Perform topology-aware scheduling and GPU resource management.
  • Oversee tenant lifecycle: onboarding, quotas, isolation, and offboarding.
  • Automate remediation workflows and runbooks; drive self-service observability.
  • Operate monitoring stack (Prometheus, Grafana, Alerts).
  • Handle incident management with customer-facing status updates and post-incident reviews.

Skills

SRE & cloud ops
GPU workload management
Kubernetes operations
Multi-tenant platforms
SLAs/SLOs management
Go or Python scripting

Tools

Terraform
Helm
GitOps (ArgoCD/Flux)
Prometheus & Grafana

Job description

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia and Ethiopia.

What you will be responsible for:

What you'll own

  • End-to-end reliability of the customer-facing GPU cloud service — availability, job completion, provisioning latency, and tenant experience.
  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs) as the runtime substrate for customer workloads.
  • Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity — placing customer jobs on the right hardware.
  • Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
  • Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
  • SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
  • Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
  • Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty — tenant-aware dashboards and alerting.
  • GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling — minimizing customer-visible impact.
  • Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
  • Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.

Customer-facing ownership

  • You are accountable for the customer's reliability experience — when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
  • Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
  • Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
  • Build self-service observability that lets customers answer their own questions — status, quota, job health — reducing support load.

Feed the AIOps substrate

  • The remediation-actuator and workflow engine land here — you make the control plane safe for automated action.
  • Your CRDs and runbooks are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter — turning customer-impacting incidents into self-healing events.

What success looks like in year 1

  • Customer-facing GPU cloud service SLAs published and met — availability, job completion, provisioning latency.
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer-visible impact.
  • BMaaS live for external tenants with self-service onboarding.
  • MTTD and MTTR for customer-impacting incidents reduced through automation.
  • Tenant self-service observability live — customers can see their own job health, quota, and status.
How you will stand out:
  • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands‑on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Customer‑facing cloud service experience — defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
  • Experience with bare‑metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation / operator development.
  • AIOps aptitude — you view the control plane as an execution surface for automated remediation, not just a scheduler.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. GPU Cloud K8S Expert (SRE SME)
Sr. GPU Cloud K8S Expert (SRE SME)

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000
Staff Slurm Cluster & HPC Engineer
Staff Slurm Cluster & HPC Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 270,000
Cloud Senior DevOps Engineer
Cloud Senior DevOps Engineer

Bitdeer • Singapore

On-site
SGD 140,000 - 210,000
Cloud Senior DevOps Engineer
Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Sr. GPU Cloud East-West Network Expert (SRE SME)
Sr. GPU Cloud East-West Network Expert (SRE SME)

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
Training and mentoring
GPU Compute & Bare Metal / DPU Engineer
GPU Compute & Bare Metal / DPU Engineer

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
GPU Compute & Bare Metal / DPU Engineer
GPU Compute & Bare Metal / DPU Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 160,000
Welfare benefits
Training and mentoring
Autonomy and growth