AI/GPU Data Center Operations Lead

Bitdeer Technologies Group

Aurora (CO)

On-site

USD 120,000 - 160,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bitdeer Technologies Group in Aurora, Colorado, is seeking a Data Center Operations Manager to lead the daily site operations and ensure availability, reliability, and operational excellence of all infrastructure and systems.

You will supervise a team of Operations Engineers, manage shift scheduling, incident response, and coordinate with engineering, network, and facilities to support AI/HPC deployments including NVIDIA B300 clusters, GPU servers, and storage.

Qualifications

  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
  • Minimum 5 years of experience in Data Center operations, IT infrastructure, or HPC/AI infrastructure management.
  • Minimum 2 years of experience in team leadership or people management.
  • Proven experience managing 24x7 shift operations in a mission-critical environment is preferred.
  • Experience with large-scale AI or HPC clusters is highly desirable.
  • Strong knowledge of Data Center operations and infrastructure management, including NVIDIA B300 Clusters, GPU Servers, x86 Servers, Storage Systems, Ethernet Networking, InfiniBand Networking.

Responsibilities

  • Lead and manage the daily operations of the Data Center site, ensuring availability, reliability, and operational excellence of all infrastructure and systems.
  • Supervise and manage a team of Operations Engineers, including manpower planning, shift scheduling, task assignment, performance management, coaching, and professional development.
  • Ensure 24x7 operational coverage and maintain adequate staffing to support business and customer requirements.
  • Act as the primary escalation point for operational incidents and coordinate cross-functional teams to drive timely issue resolution and root cause analysis.
  • Oversee the operation, maintenance, and troubleshooting of AI/HPC infrastructure, including NVIDIA B300 Clusters, GPU Servers, x86 Servers, Storage Servers, Ethernet and InfiniBand Networking, and related cabling.
  • Establish, maintain, and continuously improve operational procedures, SOPs, EOPs, and preventive maintenance programs.
  • Monitor site health, operational KPIs, incident trends, and infrastructure performance to ensure service quality and operational efficiency.
  • Coordinate hardware installation, rack and stack activities, system commissioning, infrastructure expansion, and lifecycle management.
  • Review and approve maintenance activities, change requests, incident reports, and shift handover records.
  • Ensure compliance with company policies, operational standards, safety requirements, and security procedures within the Data Center.
  • Collaborate with engineering, network, facilities, and vendor teams to support new deployments and operational improvement initiatives.
  • Participate in on‑call rotation and provide hands‑on operational support when necessary, including covering shift duties during manpower shortages, emergencies, or critical incidents.
  • Drive a culture of operational excellence, teamwork, accountability, and continuous improvement within the site operation team.

Skills

Linux administration
Incident management
Team leadership
SOP governance
Vendor coordination
Hardware troubleshooting

Education

Bachelor's degree in Computer Science / Electrical Engineering / related field

Tools

NVIDIA B300 Clusters
GPU Servers
x86 Servers
Storage Systems
Ethernet Networking
InfiniBand Networking

Job description

Bitdeer Technologies Group in Aurora, Colorado, is seeking a Data Center Operations Manager to lead the daily site operations and ensure availability, reliability, and operational excellence of all infrastructure and systems.

You will supervise a team of Operations Engineers, manage shift scheduling, incident response, and coordinate with engineering, network, and facilities to support AI/HPC deployments including NVIDIA B300 clusters, GPU servers, and storage.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Center Operations Leader - AI/HPC
Data Center Operations Leader - AI/HPC

Bitdeer (NASDAQ: BTDR) • Needham (MA)

On-site
USD 110,000 - 160,000
AI Data Center Operations Lead
AI Data Center Operations Lead

Bitdeer • Aurora (CO)

On-site
USD 120,000 - 170,000
Data Center Operations Lead for AI/HPC Infra
Data Center Operations Lead for AI/HPC Infra

Bitdeer • Needham (MA)

On-site
USD 150,000 - 210,000
Data Center Operations Lead — 24/7 AI/HPC Site
Data Center Operations Lead — 24/7 AI/HPC Site

Bitdeer (NASDAQ: BTDR) • Aurora (CO)

On-site
USD 140,000 - 180,000
AI HPC Data Center Engineer
AI HPC Data Center Engineer

Bitdeer Technologies Group • Aurora (CO)

On-site
USD 110,000 - 170,000
Data Center Operations Engineer for AI/HPC Clusters
Data Center Operations Engineer for AI/HPC Clusters

Bitdeer (NASDAQ: BTDR) • Aurora (CO)

On-site
USD 70,000 - 110,000
AI Data Center Ops Engineer - 24x7
AI Data Center Ops Engineer - 24x7

Bitdeer Technologies Group • Needham (MA)

On-site
USD 60,000 - 90,000
AI Data Center Operations Engineer
AI Data Center Operations Engineer

Bitdeer (NASDAQ: BTDR) • Needham (MA)

On-site
USD 90,000 - 150,000
AI/HPC Data Center Operations Lead
AI/HPC Data Center Operations Lead

Bitdeer Technologies Group • Needham (MA)

On-site
USD 140,000 - 210,000
AI Data Center Operations Engineer
AI Data Center Operations Engineer

Bitdeer • Needham (MA)

On-site
USD 90,000 - 130,000