Senior AI Platform Engineer

BITDEER AI PTE. LTD.

Singapore

On-site

SGD 180,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Training & mentoring
Competitive benefits
Growth opportunities

Job summary

Bitdeer Technologies Group in Singapore seeks a Senior AI Platform Engineer to keep the MaaS inference platform reliable under real customer traffic. This role owns production stability across gateway, inference-proxy, control plane, model deployments, GPU clusters, and multi-region operations.

Responsibilities include operating Kubernetes-based MaaS environments, defining SLOs, and driving safe rollouts with canaries and fast rollback.

Qualifications

  • 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services.
  • Deep Kubernetes experience, including Helm, Argo CD/GitOps, CNI/ingress, secrets, storage, and workload scheduling.
  • Hands-on experience operating GPU, AI infrastructure, or HPC workloads is strongly preferred.
  • Strong observability skills with Prometheus/VictoriaMetrics, OpenTelemetry, logs, traces, and incident diagnosis.
  • Comfortable with Go, Python, Bash, Linux networking, and production automation.
  • Proven ability to design reliable systems with clear SLOs, operational ownership, and post-incident follow-through.

Responsibilities

  • Operate and harden Kubernetes-based MaaS production environments across CPU platform nodes, edge ingress, and regional GPU tiers.
  • Define and own SLOs, alerting, dashboards, run books, and incident response for API availability, latency, error rate, capacity, and GPU health.
  • Improve rollout safety for model/runtime/platform changes using canaries, fallback, health-aware routing, maintenance mode, and fast rollback.
  • Drive capacity planning for GPU utilization, burst traffic, quota/rate limits, cross-region latency, and customer growth.
  • Automate repetitive operations through Helm, Argo CD, operators, scripts, and self-healing workflows.
  • Partner with runtime and performance engineers to debug incidents from public API edge to model worker.

Skills

Kubernetes
Helm
Argo CD
GitOps
CNI/Ingress
Observability
Prometheus
OpenTelemetry
Go
Python
Bash
Linux networking
Automation
SRE
GPU/AI infrastructure

Tools

Argo CD
Prometheus
VictoriaMetrics
OpenTelemetry
Git

Job description

About Bitdeer:

Bitdeer Technologies Group (Nasdaq: BTDR) is a world-leading technology company for Bitcoin mining. Bitdeer is committed to providing comprehensive computing solutions for its customers. The Company handles complex processes involved in computing such as equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations. The Company also offers advanced cloud capabilities to customers with high demand for artificial intelligence. Headquartered in Singapore, Bitdeer has deployed datacenters in the United States, Norway, and Bhutan.

Job Summary:

We are seeking a Senior AI Platform Engineer to keep the MaaS inference platform reliable under real customer traffic. This role owns production stability across gateway, inference-proxy, control plane, model deployments, GPU clusters, and multi-region operations.

What you will be responsible for:
  • Operate and harden Kubernetes-based MaaS production environments across CPU platform nodes, edge ingress, and regional GPU tiers.
  • Define and ownSLOs, alerting, dashboards, run books, and incident response for API availability, latency, error rate, capacity, and GPU health.
  • Improve rollout safety for model/runtime/platform changes using canaries, fallback, health-aware routing, maintenance mode, and fast rollback.
  • Drive capacity planning for GPU utilization, burst traffic, quota/rate limits, cross-region latency, and customer growth.
  • Automate repetitive operations through Helm, Argo CD, operators, scripts, and self-healing workflows.
  • Partner with runtime and performance engineers to debug incidents from public API edge to model worker.
How you will stand out:
  • 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services.
  • Deep Kubernetes experience, including Helm, Argo CD/GitOps, CNI/ingress, secrets, storage, and workload scheduling.
  • Hands-on experience operating GPU, AI infrastructure, or HPC workloads is strongly preferred.
  • Strong observability skills with Prometheus/VictoriaMetrics, OpenTelemetry, logs, traces, and incident diagnosis.
  • Comfortable with Go, Python, Bash, Linux networking, and production automation.
  • Proven ability to design reliable systems with clear SLOs, operational ownership, and post-incident follow-through.
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior MaaS Backend Engineer
Senior MaaS Backend Engineer

Bitdeer Technologies Group • Singapore

On-site
SGD 180,000 - 280,000
Senior MaaS Backend Engineer
Senior MaaS Backend Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 280,000
Cloud Senior DevOps Engineer
Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Observability & Telemetry Engineer
Senior AI Observability & Telemetry Engineer

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000
Senior Kubernetes Control Plane Engineer
Senior Kubernetes Control Plane Engineer

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
Open workspaces
Training and mentoring
Fast-growing environment
+2
SRE Monitoring Platform Software Engineer
SRE Monitoring Platform Software Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 52,000 - 76,000
Senior AI Scheduling & Orchestration Engineer
Senior AI Scheduling & Orchestration Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000