GPU Compute & Bare Metal / DPU Engineer

Bitdeer (NASDAQ: BTDR)

Penang

On-site

MYR 120,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Welfare benefits
Training and mentoring

Job summary

Bitdeer is seeking an experienced infrastructure professional to own the end-to-end lifecycle of bare-metal GPU nodes, from provisioning and onboarding through decommissioning across multiple regions. You will build automated pipelines to keep pace with fleet growth toward 10,000+ GPUs and manage firmware for DPU/SmartNIC and servers.

You will drive fleet reliability, lead incident response, and develop runbooks for on-call rotations while partnering with Storage/Image and Network teams to

Qualifications

  • 3+ years (Senior: 6+ years) in large-scale bare-metal/server-fleet operations, HPC, or cloud infrastructure.
  • Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX/DGX-class), including driver/CUDA and firmware management.
  • Strong Linux systems skills; experience with PXE/IPMI/Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU/SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking.
  • Infrastructure automation skills (Ansible, Terraform, Python/Go).
  • Comfortable owning on-call, incident management, and operational runbooks.
  • Multi-region/large-fleet operations experience is a strong plus.

Responsibilities

  • Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery/onboarding, in-service operation, break-fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU/SmartNIC and server firmware (BMC/BIOS/NIC/GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring.
  • Participate in a sustainable 7x24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage/Image and Network teams to streamline the provisioning-to-handoff path.
  • Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.

Skills

Bare-metal operations
GPU servers at scale
Linux systems
PXE/IPMI/Redfish
Automation (Ansible, Terraform, Python
On-call/incident management
Multi-region operations

Tools

Ansible
Terraform
Python
Go
PXE
IPMI
Redfish

Job description

Bitdeer is a world‑leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry‑leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About The Team

As the platform scales toward a 10,000+ GPU, multi‑region footprint, we are expanding the team that turns procured GPU capacity into reliable, sellable compute. You will own the end‑to‑end lifecycle of bare‑metal GPU nodes—from provisioning and delivery through break‑fix, firmware/DPU management, and decommissioning—directly driving delivery throughput, fleet availability, and the ROI of our largest capital investment.

What You Will Be Responsible For
  • Own the full lifecycle of bare‑metal GPU nodes: provisioning, delivery/onboarding, in‑service operation, break‑fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node‑delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU/SmartNIC and server firmware (BMC/BIOS/NIC/GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root‑cause analysis, and improve hardware‑health monitoring.
  • Participate in a sustainable 7x24 multi‑region on‑call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage/Image and Network teams to streamline the provisioning‑to‑handoff path.
  • Define bring‑up, rack, capacity, and acceptance standards for new GPU SKUs and data‑center regions.
How You Will Stand Out
  • 3+ years (Senior: 6+ years) in large‑scale bare‑metal/server‑fleet operations, HPC, or cloud infrastructure.
  • Hands‑on experience operating GPU servers at scale (e.g. NVIDIA HGX/DGX-class), including driver/CUDA and firmware management.
  • Strong Linux systems skills; experience with PXE/IPMI/Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU/SmartNIC (e.g. NVIDIA BlueField) and bare‑metal networking.
  • Infrastructure automation skills (Ansible, Terraform, Python/Go).
  • Comfortable owning on‑call, incident management, and operational runbooks.
  • Multi‑region/large‑fleet operations experience is a strong plus.
What You Will Experience Working With Us
  • A culture that values authenticity and diversity of thoughts and backgrounds.
  • An inclusive and respectable environment with open workspaces and an exciting start‑up spirit.
  • Fast‑growing company with the chance to network with industrial pioneers and enthusiasts.
  • Ability to contribute directly and make an impact on the future of the digital asset industry.
  • Involvement in new projects, developing processes/systems.
  • Personal accountability, autonomy, fast growth, and learning opportunities.
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.
Equal Employment Opportunity

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Centre Infrastructure Engineer
Data Centre Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000
Senior Security Operations Engineer, AIDC
Senior Security Operations Engineer, AIDC

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 275,157 - 353,773
Attractive welfare benefits
Personal accountability and growth opportunities
Training and mentoring programs
Cloud Senior DevOps Engineer
Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 180,000 - 300,000
Cloud Storage Engineer
Cloud Storage Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 120,000 - 240,000
AI Cloud Network Delivery Engineer
AI Cloud Network Delivery Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 120,000 - 180,000
Cloud API Engineer
Cloud API Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 180,000 - 280,000
AI Cloud Network Architect
AI Cloud Network Architect

Bitdeer (NASDAQ: BTDR) • Cyberjaya

On-site
MYR 180,000 - 320,000
AI Cloud Network Architect
AI Cloud Network Architect

Bitdeer • Malaysia

Hybrid
MYR 180,000 - 300,000
Cloud Validation & Release Engineer
Cloud Validation & Release Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 60,000 - 90,000
Data Centre Mechanical & Electrical Engineer
Data Centre Mechanical & Electrical Engineer

Bitdeer (NASDAQ: BTDR) • Cyberjaya

On-site
MYR 60,000 - 90,000
Training & mentoring