We are looking for a Data Centre Operations Engineer to support the operations and facilities of a next-generation GPU / AI data centre environment.
You will be involved in the day-to-day operation of mission-critical data centre infrastructure supporting AI and High-Performance Computing (HPC) workloads. The role combines data centre operations, facilities management, infrastructure monitoring, vendor coordination and basic server troubleshooting, with exposure to both traditional and liquid-cooled GPU infrastructure.
This is a strong opportunity for engineers looking to build deeper expertise in modern AI/HPC data centre environments.
Key Responsibilities
Data Centre Operations
- Respond to operational incidents and ensure issues are resolved or escalated according to severity, impact and SLA requirements.
- Perform hands-on operational activities involving electrical, air-cooling and liquid-cooling systems.
- Monitor the physical health and operating conditions of GPU and data centre infrastructure.
- Conduct visual inspections of servers, Cooling Distribution Units (CDUs) and related equipment.
- Support server troubleshooting together with remote engineering and technical teams.
- Coordinate access and security clearance for vendors and visitors.
- Ensure vendors comply with workplace safety, security and data centre operating requirements.
- Contribute to the continuous improvement of operational processes and procedures for GPU-oriented environments.
Data Centre Facilities
- Monitor critical facilities infrastructure, including:Power and electrical systemsAir and liquid cooling systemsLeakage detectionEnvironmental monitoring and controlsBuilding Management Systems (BMS)
- Coordinate preventive maintenance, planned shutdowns and infrastructure works with internal stakeholders and external vendors.
- Ensure adherence to Standard Operating Procedures (SOPs), Methods of Procedure (MOPs) and Emergency Response Procedures (ERPs).
- Maintain accurate data centre documentation and prepare operational and facilities reports.
- Prepare monthly facilities management reports covering data centre health and operational status.
- Identify potential workplace safety, operational and infrastructure risks.
- Support capacity and operational planning by applying knowledge of power and cooling requirements for high-density GPU infrastructure.
- Work with multiple technical teams and stakeholders to resolve operational and facilities-related issues.
Requirements
- Diploma or higher qualification in Mechanical Engineering, Electrical Engineering, Building Services or a related discipline.
- Good understanding of mission-critical data centre infrastructure, particularly:Electrical and mechanical systemsPower and cooling infrastructureFire protection and safety systemsBuilding Management Systems (BMS)Equipment maintenance
- Experience supporting the maintenance and operation of data centre electrical and/or mechanical infrastructure.
- Exposure to liquid cooling, high-density computing, GPU infrastructure or AI/HPC environments would be advantageous, but is not essential.
- Comfortable working in a hands-on, operational data centre environment.
- Able to work independently while collaborating effectively with technical teams, stakeholders and vendors.
- Organised, adaptable and comfortable responding to changing operational requirements.
- Strong willingness to learn emerging GPU, AI/HPC and next-generation data centre technologies.
- Willing to provide support outside standard business hours when required, including nights, weekends and public holidays.