GPU Compute & Bare Metal / DPU Engineer

Bitdeer (NASDAQ: BTDR)

San Jose (CA)

On-site

USD 170,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Bitdeer Technologies Group in the San Jose area is seeking a seasoned professional to own the end-to-end lifecycle of bare-metal GPU nodes, from provisioning to decommissioning, across multiple regions, driving delivery throughput and ROI on our GPU fleet.

You will build automated pipelines, manage DPU/SmartNIC and firmware, improve MTTR, lead on-call rotations, and collaborate with Storage, Image and Network teams to define bring-up standards for new GPU SKUs and data-center regions.

Qualifications

  • 3+ years in large-scale bare-metal/server-fleet ops, HPC, or cloud infrastructure.
  • Hands-on GPU server experience at scale (NVIDIA HGX/DGX-class) with driver/CUDA and firmware.
  • Strong Linux systems skills; experience with PXE/IPMI/Redfish, OS imaging, and automated provisioning.

Responsibilities

  • Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery/onboarding, in-service operation, break-fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring.
  • Participate in a sustainable 7x24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage / Image and Network teams to streamline the provisioning-to-handoff path.
  • Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.

Skills

Linux systems
PXE/IPMI/Redfish
OS imaging
Infrastructure automation
Ansible
Terraform
Python/Go
On-call/Incidents
Multi-region operations
NVIDIA BlueField

Tools

NVIDIA HGX/DGX-class

Job description

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

About Bitdeer Technologies Group

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

As the platform scales toward a 10,000+ GPU, multi-region footprint, we are expanding the team that turns procured GPU capacity into reliable, sellable compute. You will own the end-to-end lifecycle of bare-metal GPU nodes - provisioning and delivery through break-fix, firmware/DPU management, and decommissioning - directly driving delivery throughput, fleet availability, and the ROI of our largest capital investment.

Key Responsibilities
  • Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery / onboarding, in-service operation, break-fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring.
  • Participate in a sustainable 7x24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage / Image and Network teams to streamline the provisioning-to-handoff path.
  • Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.
Job Requirement
  • 3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations, HPC, or cloud infrastructure.
  • Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX-class), including driver / CUDA and firmware management.
  • Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking.
  • Infrastructure automation skills (Ansible, Terraform, Python / Go).
  • Comfortable owning on-call, incident management, and operational runbooks.
  • Multi-region / large-fleet operations experience is a strong plus.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Compute & Bare Metal / DPU Engineer
GPU Compute & Bare Metal / DPU Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 120,000 - 180,000
GPU Compute & Bare Metal / DPU Engineer
GPU Compute & Bare Metal / DPU Engineer

Bitdeer • San Jose (CA)

On-site
USD 150,000 - 210,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer Technologies Group • Austin (TX)

On-site
USD 170,000 - 250,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 180,000 - 240,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer Technologies Group • San Jose (CA)

On-site
USD 180,000 - 240,000
GPU Bare-Metal & DPU Engineer for Scalable Compute
GPU Bare-Metal & DPU Engineer for Scalable Compute

Bitdeer • San Jose (CA)

On-site
USD 150,000 - 210,000
GPU Bare-Metal & DPU Engineer — Automation & Ops
GPU Bare-Metal & DPU Engineer — Automation & Ops

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 170,000 - 250,000
Cloud Network / SDN Engineer
Cloud Network / SDN Engineer

Bitdeer (NASDAQ: BTDR) • San Jose (CA)

On-site
USD 140,000 - 220,000
Cloud Network / SDN Engineer
Cloud Network / SDN Engineer

Bitdeer • San Jose (CA)

On-site
USD 140,000 - 200,000