GPU Compute & Bare Metal / DPU Engineer

Bitdeer Group

Singapore

On-site

SGD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Welfare benefits
Training and mentoring
Autonomy and growth

Job summary

Bitdeer Group is seeking an experienced professional to own the end-to-end lifecycle of bare-metal GPU nodes, from provisioning to decommissioning, across multiple regions. You will build automated pipelines to keep pace with fleet growth toward 10,000+ GPUs and manage firmware across DPU/SmartNIC and servers.

You will operate Linux-based environments, apply automation (Ansible, Terraform, Python/Go) and participate in a 24/7 on-call rotation, delivering reliable compute for a growing cloud and

Qualifications

  • 3+ years in large-scale bare-metal / server-fleet operations, HPC, or cloud infra.
  • Hands-on experience operating GPU servers at scale (NVIDIA HGX/DGX).
  • Strong Linux skills; familiarity with PXE/IPMI/Redfish and automated provisioning.
  • Experience with DPU/SmartNIC and bare-metal networking is a plus.
  • Automation via Ansible, Terraform, Python/Go; on-call incident management.

Responsibilities

  • Own full lifecycle of bare-metal GPU nodes across regions: provisioning, delivery, in-service operation, break-fix, decommissioning.
  • Build and operate automated node-delivery pipelines to reduce backlog and scale fleet.
  • Manage DPU/SmartNIC and server firmware baselines, upgrades, validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and RCA, improve hardware health monitoring.
  • Participate in 24/7 multi-region on-call rotation; develop runbooks and tooling to reduce toil.
  • Collaborate with Storage/Image and Network teams to streamline provisioning-to-handoff.
  • Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-centre regions.

Skills

Bare-metal server operations
GPU servers at scale
Linux systems
PXE
IPMI
Redfish
Ansible
Terraform
Python
Go
On-call incident management
Multi-region operations

Tools

NVIDIA BlueField
BMC firmware management

Job description

Bitdeer is a world‑leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry‑leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About the team :

As the platform scales toward a 10,000+ GPU, multi‑region footprint, we are expanding the team that turns procured GPU capacity into reliable, sellable compute. You will own the end‑to‑end lifecycle of bare‑metal GPU nodes—from provisioning and delivery through break‑fix, firmware/DPU management, and decommissioning—directly driving delivery throughput, fleet availability, and the ROI of our largest capital investment.

What you will be responsible for:
  • Own the full lifecycle of bare‑metal GPU nodes: provisioning, delivery / onboarding, in‑service operation, break‑fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node‑delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root‑cause analysis, and improve hardware‑health monitoring.
  • Participate in a sustainable 24/7 multi‑region on‑call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage / Image and Network teams to streamline the provisioning‑to‑handoff path.
  • Define bring‑up, rack, capacity, and acceptance standards for new GPU SKUs and data‑center regions.
How you will stand out:
  • 3+ years (Senior: 6+ years) in large‑scale bare‑metal / server‑fleet operations, HPC, or cloud infrastructure.
  • Hands‑on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX‑class), including driver / CUDA and firmware management.
  • Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare‑metal networking.
  • Infrastructure automation skills (Ansible, Terraform, Python / Go).
  • Comfortable owning on‑call, incident management, and operational runbooks.
  • Multi‑region / large‑fleet operations experience is a strong plus.
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds.
  • An inclusive and respectable environment with open workspaces and exciting start‑up spirit.
  • Fast‑growing company with the chance to network with industrial pioneers and enthusiasts.
  • Ability to contribute directly and make an impact on the future of the digital asset industry.
  • Involvement in new projects, developing processes/systems.
  • Personal accountability, autonomy, fast growth, and learning opportunities.
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Compute & Bare Metal / DPU Engineer
GPU Compute & Bare Metal / DPU Engineer

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
SRE L1 Support/Cloud Platform Ops Engineers
SRE L1 Support/Cloud Platform Ops Engineers

United States Digital Space LLC • Singapore

On-site
SGD 60,000 - 90,000
SRE L1 Support/Cloud Platform Ops Engineers
SRE L1 Support/Cloud Platform Ops Engineers

Bitdeer Technologies Group • Singapore

On-site
SGD 4,000 - 7,000
Welfare benefits
Training & mentoring
Inclusive culture
SRE L1 Support/Cloud Platform Ops Engineers
SRE L1 Support/Cloud Platform Ops Engineers

Bitdeer • Singapore

On-site
SGD 52,000 - 78,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Cloud Network Operations Engineer
Senior AI Cloud Network Operations Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 160,000
Attractive welfare benefits
Developmental opportunities
Flexible working environment
Senior GPU Bare-Metal Engineer | DPU & HPC Automation
Senior GPU Bare-Metal Engineer | DPU & HPC Automation

Bitdeer Group • Singapore

On-site
SGD 120,000 - 160,000
Welfare benefits
Training and mentoring
Autonomy and growth
Cloud Storage Engineer
Cloud Storage Engineer

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000