AI/HPC Data Center Operations Engineer

Bitdeer Technologies Group

Cyberjaya

On-site

MYR 60,000 - 100,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bitdeer Technologies Group is seeking a Data Center Operations Engineer to manage day-to-day data center infrastructure, including installation, maintenance, and troubleshooting of AI/HPC clusters (GB200/GB300) and related GPU/x86/storage servers. You will monitor health, perform firmware updates and diagnostics, and support provisioning and network validation.

The role requires Linux admin skills, TCP/IP networking knowledge, and a proactive, detail-oriented approach.

Qualifications

  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
  • Basic understanding of Data Center infrastructure and server hardware architecture.
  • Familiarity with one or more of the following systems: NVIDIA GB200 Cluster, NVIDIA GB300 Cluster, GPU Servers, x86 Servers, Storage Servers, Ethernet and InfiniBand Networks.
  • Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
  • Familiarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.
  • Understanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.
  • Basic Linux administration skills, including: System monitoring and troubleshooting, Service management using systemctl, Log analysis using journalctl and dmesg, Network troubleshooting tools such as ip and ethtool, Basic shell scripting.
  • Preferred Qualifications: Experience in Data Center operations or hardware maintenance; Experience supporting AI/HPC infrastructure or GPU clusters; Familiarity with NVIDIA AI infrastructure GB200/GB300; Experience with large-scale cluster environments and high-speed networking technologies; Familiarity with monitoring/orchestration tools such as Slurm, Kubernetes, Prometheus, Grafana.
  • Personal Attributes: Willingness to work in a 24x7 two-shift rotation schedule, strong sense of responsibility, teamwork and communication, ability to work under pressure, detail-oriented, self-motivated.

Responsibilities

  • Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
  • Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including NVIDIA GB200 and GB300 clusters, GPU Servers, x86 Servers, Storage Servers, and related networking gear.
  • Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
  • Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
  • Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
  • Troubleshoot hardware and infrastructure issues including server failures, GPU errors, storage issues, network problems, switch failures, and cabling faults.
  • Perform routine inspections, preventive maintenance, and maintain accurate operation logs and SOPs.
  • Execute incident response procedures and escalation per standards; prepare shift handover reports.
  • Collaborate with engineering, network, and infrastructure teams to support new deployments and improvements.
  • Participate in a two-shift rotation schedule, including night shifts, weekends, and holidays as required.

Skills

Data Center Infra
GB200 cluster
GB300 cluster
GPU Servers
x86 Servers
Storage Servers
InfiniBand / RoCE
TCP/IP Networking
Linux Administration
Shell Scripting
Monitoring Tools
Hardware Diagnostics

Education

Bachelor's degree in Computer Science/Engineering or related

Tools

Slurm
Kubernetes
Prometheus
Grafana

Job description

Bitdeer Technologies Group is seeking a Data Center Operations Engineer to manage day-to-day data center infrastructure, including installation, maintenance, and troubleshooting of AI/HPC clusters (GB200/GB300) and related GPU/x86/storage servers. You will monitor health, perform firmware updates and diagnostics, and support provisioning and network validation.

The role requires Linux admin skills, TCP/IP networking knowledge, and a proactive, detail-oriented approach.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Operations Engineer - AI/HPC GPU Clusters
Data Center Operations Engineer - AI/HPC GPU Clusters

Bitdeer (NASDAQ: BTDR) • Johor Bahru

On-site
MYR 100,000 - 145,000
Open workspaces
Training and mentoring
Welfare benefits
AI/HPC Data Center Infrastructure Engineer
AI/HPC Data Center Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000
AI/HPC Data Center Operations Engineer
AI/HPC Data Center Operations Engineer

Bitdeer • Cyberjaya

On-site
MYR 67,000 - 100,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer • Cyberjaya

On-site
MYR 67,000 - 100,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer Technologies Group • Cyberjaya

On-site
MYR 60,000 - 100,000
Data Centre Infrastructure Engineer
Data Centre Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer (NASDAQ: BTDR) • Johor Bahru

On-site
MYR 100,000 - 145,000
Open workspaces
Training and mentoring
Welfare benefits
GPU Bare-Metal & DPU Engineer for Scalable HPC Fleet
GPU Bare-Metal & DPU Engineer for Scalable HPC Fleet

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 120,000 - 190,000
Welfare benefits
Training and mentoring
Senior AI Platform Engineer - Kubernetes & GPU Infra Lead
Senior AI Platform Engineer - Kubernetes & GPU Infra Lead

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 180,000 - 360,000
Senior Data Center Ops Engineer — GPU & AI Infra
Senior Data Center Ops Engineer — GPU & AI Infra

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000