AI Data Center Operations Lead

Bitdeer

Aurora (CO)

On-site

USD 120,000 - 170,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bitdeer is seeking a competent Data Center Operations Lead to manage daily site activities and ensure uptime for AI/HPC infrastructure. You will supervise Operations Engineers, plan staffing, and act as the primary escalation point for incidents, coordinating cross-functional teams to resolve issues quickly.

The role requires 5+ years in data center or HPC operations, and 2+ years in leadership. You will work with NVIDIA GPU clusters, GPU servers, and networked storage, while driving SOPs and

Qualifications

  • Bachelor’s degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
  • Minimum 5 years of experience in Data Center operations, IT infrastructure, or HPC/AI infrastructure management.
  • Minimum 2 years of experience in team leadership or people management.
  • Proven experience managing 24x7 shift operations in a mission-critical environment is preferred.
  • Experience with large‑scale AI or HPC clusters is highly desirable.
  • Strong knowledge of Data Center operations and infrastructure management, including NVIDIA B300 Clusters, GPU Servers, x86 Servers, Storage Systems, Ethernet Networking, InfiniBand Networking.
  • Familiarity with NVIDIA GPU architecture, NVLink, NVSwitch, and AI cluster deployment concepts.
  • Experience with server hardware troubleshooting, firmware management, and hardware lifecycle management.
  • Good understanding of structured cabling systems, including optical fiber, MPO/LC connectors, DAC, and AOC cabling.
  • Familiarity with infrastructure monitoring and management tools.
  • Strong Linux administration and troubleshooting skills, including System and service management, Hardware and performance diagnostics, Log analysis and incident investigation, Network troubleshooting, Basic scripting and automation.
  • Demonstrated ability to lead, motivate, and develop a team of Operations Engineers.
  • Shift scheduling and workforce planning, Incident and escalation management, Performance management and coaching, SOP/EOP development and operational governance, Vendor and stakeholder coordination.
  • Willingness to provide hands-on operational support and participate in on-call duties when required.
  • Strong ownership mindset and accountability for site operations.
  • Excellent communication and interpersonal skills.
  • Ability to remain calm and make sound decisions during critical incidents.
  • Highly organized, detail-oriented, and committed to operational excellence and continuous improvement.

Responsibilities

  • Lead and manage the daily operations of the Data Center site, ensuring availability, reliability, and operational excellence of infrastructure and systems.
  • Supervise and manage a team of Operations Engineers, including manpower planning, shift scheduling, task assignment, performance management, coaching, and professional development.
  • Ensure 24x7 operational coverage and maintain adequate staffing to support business and customer requirements.
  • Act as the primary escalation point for operational incidents and coordinate cross-functional teams to drive timely issue resolution and root cause analysis.
  • Oversee operation, maintenance, and troubleshooting of AI/HPC infrastructure, including NVIDIA B300 Clusters, GPU Servers, x86 Servers, Storage Servers, Ethernet and InfiniBand Switches, and cabling systems.
  • Establish, maintain, and continuously improve SOPs, EOPs, and preventive maintenance programs.
  • Monitor site health, KPIs, incident trends, and infrastructure performance to ensure service quality.
  • Coordinate hardware installation, rack/stack, system commissioning, infrastructure expansion, and lifecycle management.
  • Review and approve maintenance activities, change requests, incident reports, and shift handover records.
  • Ensure compliance with policies, standards, safety requirements, and security procedures within the Data Center.
  • Collaborate with engineering, network, facilities, and vendor teams to support deployments and improvements.
  • Participate in on-call rotation and provide hands-on operational support when necessary.
  • Drive a culture of operational excellence, teamwork, accountability, and continuous improvement.

Skills

Linux administration
Leadership
Incident management
On-call readiness
Scripting & automation

Education

Bachelor's degree in Computer Science/Engineering or related

Tools

NVIDIA GPU hardware
NVLink/NVSwitch
Storage systems
Ethernet networking
InfiniBand networking
Server firmware management

Job description

Bitdeer is seeking a competent Data Center Operations Lead to manage daily site activities and ensure uptime for AI/HPC infrastructure. You will supervise Operations Engineers, plan staffing, and act as the primary escalation point for incidents, coordinating cross-functional teams to resolve issues quickly.

The role requires 5+ years in data center or HPC operations, and 2+ years in leadership. You will work with NVIDIA GPU clusters, GPU servers, and networked storage, while driving SOPs and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/GPU Data Center Operations Lead
AI/GPU Data Center Operations Lead

Bitdeer Technologies Group • Aurora (CO)

On-site
USD 120,000 - 160,000
Data Center Operations Leader - AI/HPC
Data Center Operations Leader - AI/HPC

Bitdeer (NASDAQ: BTDR) • Needham (MA)

On-site
USD 110,000 - 160,000
Data Center Operations Lead for AI/HPC Infra
Data Center Operations Lead for AI/HPC Infra

Bitdeer • Needham (MA)

On-site
USD 150,000 - 210,000
Data Center Operations Lead — 24/7 AI/HPC Site
Data Center Operations Lead — 24/7 AI/HPC Site

Bitdeer (NASDAQ: BTDR) • Aurora (CO)

On-site
USD 140,000 - 180,000
AI Data Center Operations Engineer
AI Data Center Operations Engineer

Bitdeer (NASDAQ: BTDR) • Needham (MA)

On-site
USD 90,000 - 150,000
AI Data Center Ops Engineer - 24x7
AI Data Center Ops Engineer - 24x7

Bitdeer Technologies Group • Needham (MA)

On-site
USD 60,000 - 90,000
AI/HPC Data Center Operations Lead
AI/HPC Data Center Operations Lead

Bitdeer Technologies Group • Needham (MA)

On-site
USD 140,000 - 210,000
AI Data Center Operations Engineer — 24/7 Shift
AI Data Center Operations Engineer — 24/7 Shift

Bitdeer • South Dakota

On-site
USD 70,000 - 90,000
Data Center Operations Engineer for AI/HPC Clusters
Data Center Operations Engineer for AI/HPC Clusters

Bitdeer (NASDAQ: BTDR) • Aurora (CO)

On-site
USD 70,000 - 110,000
AI Data Center Operations Engineer
AI Data Center Operations Engineer

Bitdeer • Needham (MA)

On-site
USD 90,000 - 130,000