Senior AI Platform Engineer — Kubernetes & GPU Infra

Bitdeer Group

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bitdeer Group in Singapore is seeking a Senior Site Reliability Engineer to own and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers.

You will define SLOs, incident response, automate deployments with Helm and Argo CD, and collaborate with runtime and performance engineers to debug incidents from edge to model worker. This role offers growth in a fast-growing cloud platform for digital assets.

Qualifications

  • 6+ years in SRE, platform engineering, or infra for production cloud services.
  • Deep Kubernetes experience with Helm, Argo CD/GitOps, CNI/Ingress, secrets, storage, scheduling.
  • Hands-on GPU/AI infrastructure or HPC workload experience preferred.
  • Strong observability with Prometheus/OpenTelemetry, logs and traces.
  • Go, Python, Bash, Linux networking, and production automation.
  • Ability to design reliable systems with clear SLOs and incident follow-through.

Responsibilities

  • Operate and harden Kubernetes-based MaaS environments across CPU nodes, edge ingress, and regional GPU tiers.
  • Define and own SLOs, alerts, dashboards, runbooks, and incident response for API and GPU health.
  • Improve rollout safety for changes using canaries, fallback, and quick rollback.
  • Drive capacity planning for GPU utilization, burst traffic, quotas, and cross-region latency.
  • Automate operations with Helm, Argo CD, operators, and self-healing workflows.
  • Collaborate with runtime/perf engineers to debug incidents from edge to model worker.

Skills

Kubernetes
SRE/Platform
Prometheus/OpenTelemetry
Go/Python scripting
Argo CD
GPU/AI infra

Tools

Helm
Argo CD
GitOps
CI/CD

Job description

Bitdeer Group in Singapore is seeking a Senior Site Reliability Engineer to own and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers.

You will define SLOs, incident response, automate deployments with Helm and Argo CD, and collaborate with runtime and performance engineers to debug incidents from edge to model worker. This role offers growth in a fast-growing cloud platform for digital assets.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer - Kubernetes & GPU Cloud
Senior AI Platform Engineer - Kubernetes & GPU Cloud

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
Senior GPU Systems & Fabric Architect
Senior GPU Systems & Fabric Architect

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000
Senior Kubernetes Control Plane Architect
Senior Kubernetes Control Plane Architect

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Scheduling & HPC Orchestration Engineer
Senior AI Scheduling & HPC Orchestration Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 170,000 - 240,000
Senior AI Telemetry & Observability Engineer
Senior AI Telemetry & Observability Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 240,000
Senior Cloud DevOps & MLOps Engineer
Senior Cloud DevOps & MLOps Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
AI Platform Deployment Engineer
AI Platform Deployment Engineer

GTS Consulting • Singapore

On-site
SGD 100,000 - 180,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000