Data Center Operations Engineer

Bitdeer Technologies Group

Aurora (CO)

On-site

USD 110,000 - 170,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bitdeer Technologies Group in Colorado is seeking a Data Center Operations Technician to manage daily operation and maintenance of AI/HPC cluster infrastructure, including NVIDIA B300 systems, GPU servers, x86 servers, storage, and networking gear. You will handle hardware provisioning, BIOS/BMC updates, fault diagnosis, incident response, and shift handovers in a 24x7 rotation.

A bachelor's degree in a relevant field and strong Linux basics are required, with a passion for reliability and

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related discipline.
  • Basic Linux administration and monitoring.
  • Familiarity with data center hardware and networking concepts.
  • Experience with AI/HPC infrastructure or GPU clusters is a plus.

Responsibilities

  • Manage daily operation and maintenance of Data Center infrastructure for high availability.
  • Install, rack and stack, cabling, commissioning, maintenance and troubleshooting of AI/HPC cluster infrastructure.
  • Monitor health of cluster systems including servers, GPUs, storage and networking devices.
  • Perform hardware replacement, BIOS/BMC/Firmware upgrades, and diagnostics.
  • Support server provisioning, OS installation, cluster expansion, network validation, and burn-in testing.
  • Troubleshoot hardware/infrastructure issues and log incidents.
  • Maintain operation logs and SOPs; prepare shift handover reports.
  • Collaborate with engineering, network and infra teams for deployments and improvements.
  • Participate in a 3-shift rotation including nights, weekends, holidays.

Skills

Data Center operations
GPU clusters
Linux administration
NVIDIA B300 knowledge
Networking basics

Education

Bachelor's degree in CS/CE/EE or related

Tools

NVIDIA B300 Cluster
GPU Servers
x86 Servers
Storage Servers
InfiniBand Networking

Job description

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Key Responsibilities
  1. Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
  2. Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including:
    • NVIDIA B300 Cluster
    • GPU Servers
    • x86 Servers
    • Storage Servers
    • Ethernet and InfiniBand Switches
    • DAC, AOC, Optical Fiber, and related cabling infrastructure
  3. Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
  4. Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
  5. Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
  6. Troubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.
  7. Perform routine inspections, preventive maintenance, and maintain accurate operational records and maintenance logs.
  8. Execute incident response procedures and provide timely escalation and resolution according to operational standards.
  9. Prepare shift handover reports and maintain operation documents, SOPs, and incident reports.
  10. Work closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.
  11. Participate in a three-shift rotation schedule, including night shifts, weekends, and holidays as required.
Requirements
Education
  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
Technical Skills
  • Basic understanding of Data Center infrastructure and server hardware architecture.
  • Familiarity with one or more of the following systems:
    • NVIDIA B300 Cluster
    • GPU Servers
    • x86 Servers
    • Storage Servers
    • Ethernet and InfiniBand Networks
  • Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
  • Familiarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.
  • Understanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.
Linux Skills
  • Basic Linux administration skills, including:
    • System monitoring and troubleshooting
    • Service management using systemctl
    • Log analysis using journalctl and dmesg
    • Network troubleshooting tools such as ip and ethtool
    • Basic shell scripting
Preferred Qualifications
  • Experience in Data Center operations or hardware maintenance is preferred.
  • Experience supporting AI/HPC infrastructure or GPU clusters is a plus.
  • Familiarity with NVIDIA AI infrastructure, including GB200 and GB300 systems, is highly desirable.
  • Experience with large-scale cluster environments and high-speed networking technologies is a plus.
  • Familiarity with monitoring and orchestration tools such as Slurm, Kubernetes, Prometheus, or Grafana is an advantage.
Personal Attributes
  • Willingness to work in a 24x7 shift rotation schedule, including night shifts.
  • Strong sense of responsibility and ownership.
  • Good teamwork and communication skills.
  • Ability to work under pressure and respond effectively to operational incidents.
  • Detail-oriented with strong adherence to operational procedures and safety standards.
  • Self-motivated with a proactive attitude toward learning and problem-solving.

This position is ideal for candidates who are interested in building and operating next-generation AI Data Center infrastructure supporting large-scale NVIDIA B300 clusters.

Equal Opportunity Employer

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer • Needham (MA)

On-site
USD 90,000 - 130,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer • South Dakota

On-site
USD 70,000 - 90,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer Technologies Group • Needham (MA)

On-site
USD 60,000 - 90,000
Data Center Site Manager / Supervisor
Data Center Site Manager / Supervisor

Bitdeer • Aurora (CO)

On-site
USD 120,000 - 170,000
Data Center Site Manager / Supervisor
Data Center Site Manager / Supervisor

Bitdeer • Needham (MA)

On-site
USD 150,000 - 210,000
Data Center Site Manager / Supervisor
Data Center Site Manager / Supervisor

Bitdeer Technologies Group • Needham (MA)

On-site
USD 140,000 - 210,000
Data Center Site Manager / Supervisor
Data Center Site Manager / Supervisor

Bitdeer Technologies Group • Aurora (CO)

On-site
USD 120,000 - 160,000
Senior AI Data Center Network Engineer
Senior AI Data Center Network Engineer

Bitdeer • Aurora (CO)

On-site
USD 140,000 - 210,000
Senior AI Data Center Network Engineer
Senior AI Data Center Network Engineer

Bitdeer Technologies Group • Needham (MA)

On-site
USD 180,000 - 240,000
Senior AI Data Center Network Engineer
Senior AI Data Center Network Engineer

Bitdeer • Needham (MA)

On-site
USD 150,000 - 210,000