Data Center Operations Engineer

Bitdeer Technologies Group

Cyberjaya

On-site

MYR 60,000 - 100,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer Technologies Group is seeking a Data Center Operations Engineer to manage day-to-day data center infrastructure, including installation, maintenance, and troubleshooting of AI/HPC clusters (GB200/GB300) and related GPU/x86/storage servers. You will monitor health, perform firmware updates and diagnostics, and support provisioning and network validation.

The role requires Linux admin skills, TCP/IP networking knowledge, and a proactive, detail-oriented approach.

Qualifications

  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
  • Basic understanding of Data Center infrastructure and server hardware architecture.
  • Familiarity with one or more of the following systems: NVIDIA GB200 Cluster, NVIDIA GB300 Cluster, GPU Servers, x86 Servers, Storage Servers, Ethernet and InfiniBand Networks.
  • Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
  • Familiarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.
  • Understanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.
  • Basic Linux administration skills, including: System monitoring and troubleshooting, Service management using systemctl, Log analysis using journalctl and dmesg, Network troubleshooting tools such as ip and ethtool, Basic shell scripting.
  • Preferred Qualifications: Experience in Data Center operations or hardware maintenance; Experience supporting AI/HPC infrastructure or GPU clusters; Familiarity with NVIDIA AI infrastructure GB200/GB300; Experience with large-scale cluster environments and high-speed networking technologies; Familiarity with monitoring/orchestration tools such as Slurm, Kubernetes, Prometheus, Grafana.
  • Personal Attributes: Willingness to work in a 24x7 two-shift rotation schedule, strong sense of responsibility, teamwork and communication, ability to work under pressure, detail-oriented, self-motivated.

Responsibilities

  • Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
  • Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including NVIDIA GB200 and GB300 clusters, GPU Servers, x86 Servers, Storage Servers, and related networking gear.
  • Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
  • Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
  • Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
  • Troubleshoot hardware and infrastructure issues including server failures, GPU errors, storage issues, network problems, switch failures, and cabling faults.
  • Perform routine inspections, preventive maintenance, and maintain accurate operation logs and SOPs.
  • Execute incident response procedures and escalation per standards; prepare shift handover reports.
  • Collaborate with engineering, network, and infrastructure teams to support new deployments and improvements.
  • Participate in a two-shift rotation schedule, including night shifts, weekends, and holidays as required.

Skills

Data Center Infra
GB200 cluster
GB300 cluster
GPU Servers
x86 Servers
Storage Servers
InfiniBand / RoCE
TCP/IP Networking
Linux Administration
Shell Scripting
Monitoring Tools
Hardware Diagnostics

Education

Bachelor's degree in Computer Science/Engineering or related

Tools

Slurm
Kubernetes
Prometheus
Grafana

Job description

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

What you will be responsible for:
  • Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
  • Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including:
    • NVIDIA GB200 Cluster
    • NVIDIA GB300 Cluster
    • GPU Servers
    • x86 Servers
    • Storage Servers
    • Ethernet and InfiniBand Switches
    • DAC, AOC, Optical Fiber, and related cabling infrastructure
  • Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
  • Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
  • Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
  • Troubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.
  • Perform routine inspections, preventive maintenance, and maintain accurate operational records and maintenance logs.
  • Execute incident response procedures and provide timely escalation and resolution according to operational standards.
  • Prepare shift handover reports and maintain operation documents, SOPs, and incident reports.
  • Work closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.
  • Participate in a two-shift rotation schedule, including night shifts, weekends, and holidays as required.
How you will stand out:
  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
  • Basic understanding of Data Center infrastructure and server hardware architecture.
  • Familiarity with one or more of the following systems:
    • NVIDIA GB200 Cluster
    • NVIDIA GB300 Cluster
    • GPU Servers
    • x86 Servers
    • Storage Servers
    • Ethernet and InfiniBand Networks
  • Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
  • Familiarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.
  • Understanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.
  • Basic Linux administration skills, including:
    • System monitoring and troubleshooting
    • Service management using systemctl
    • Log analysis using journalctl and dmesg
    • Network troubleshooting tools such as ip and ethtool
    • Basic shell scripting
  • Preferred Qualifications
    • Experience in Data Center operations or hardware maintenance is preferred.
    • Experience supporting AI/HPC infrastructure or GPU clusters is a plus.
    • Familiarity with NVIDIA AI infrastructure, including GB200 and GB300 systems, is highly desirable.
    • Experience with large-scale cluster environments and high-speed networking technologies is a plus.
    • Familiarity with monitoring and orchestration tools such as Slurm, Kubernetes, Prometheus, or Grafana is an advantage.
  • Personal Attributes
    • Willingness to work in a 24x7 two-shift rotation schedule, including night shifts.
    • Strong sense of responsibility and ownership.
    • Good teamwork and communication skills.
    • Ability to work under pressure and respond effectively to operational incidents.
    • Detail-oriented with strong adherence to operational procedures and safety standards.
    • Self-motivated with a proactive attitude toward learning and problem-solving.
    • This position is ideal for candidates who are interested in building and operating next-generation AI Data Center infrastructure supporting Large-scale NVIDIA GB200 and GB300 GPU clusters.
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Centre Infrastructure Engineer
Data Centre Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000
AI Cloud Network Delivery Engineer
AI Cloud Network Delivery Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 120,000 - 180,000
AI Cloud Network Architect
AI Cloud Network Architect

Bitdeer (NASDAQ: BTDR) • Cyberjaya

On-site
MYR 180,000 - 320,000
Cloud Senior DevOps Engineer
Cloud Senior DevOps Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 180,000 - 300,000
Senior Security Operations Engineer, AIDC
Senior Security Operations Engineer, AIDC

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 275,157 - 353,773
Attractive welfare benefits
Personal accountability and growth opportunities
Training and mentoring programs
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 180,000 - 360,000
AI Cloud Network Architect
AI Cloud Network Architect

Bitdeer • Malaysia

Hybrid
MYR 180,000 - 300,000
Data Centre Mechanical & Electrical Engineer
Data Centre Mechanical & Electrical Engineer

Bitdeer (NASDAQ: BTDR) • Cyberjaya

On-site
MYR 60,000 - 90,000
Training & mentoring
AI Cloud Network Architect
AI Cloud Network Architect

Bitdeer Technologies Group • Cyberjaya

Hybrid
MYR 240,000 - 360,000
Data Centre Mechanical & Electrical Engineer
Data Centre Mechanical & Electrical Engineer

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 67,000 - 100,000