Reliability Engineer - Mechanical

Khazna Data Centers

Dubai

On-site

AED 180,000 - 300,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Khazna Data Centers in Dubai seeks a Site Reliability Engineer to advance reliability programs across multiple data centers. You will monitor power, cooling and IT systems, drive preventive maintenance, and lead root cause analyses to minimize downtime.

Reporting to the Reliability Manager, you will collaborate with Operations, Engineering, Facilities and vendors to implement availability plans, track KPIs (MTBF, MTTR, uptime) and develop predictive maintenance routines using IoT, data analytics

Qualifications

  • Bachelor’s degree in mechanical, electrical, reliability, or related Engineering discipline.
  • 3+ years of experience in reliability engineering, maintenance engineering, or data center operations.
  • Hands‑on experience with RCA, FMEA, and predictive maintenance methodologies.
  • Proficiency with monitoring platforms, data‑analytics tools, and scripting (e.g., Python, R).
  • Familiarity with IoT sensors, machine‑learning frameworks, and condition-based monitoring systems.
  • Knowledge of industry reliability standards and regulations (ISO, ASHRAE, Uptime Institute).

Responsibilities

  • Monitor real-time and historical performance metrics for critical power, cooling, and IT systems.
  • Analyse system data to identify trends, failure modes, and reliability risks.
  • Execute RCA and FMEA, then drive corrective and preventive actions.
  • Develop and maintain condition-based and predictive maintenance routines using IoT, data analytics, and ML tools.
  • Support preventive maintenance programs: schedule, document, and validate maintenance activities.
  • Assist in asset lifecycle planning, including upgrades and end-of-life strategies.
  • Contribute to capacity runway assessments to forecast infrastructure needs.
  • Implement and enforce availability management plans, risk assessments, and mitigation strategies.
  • Ensure data collection and reporting for reliability KPIs (MTBF, MTTR, availability).
  • Prepare reliability reports and dashboards; present findings to site leadership.
  • Respond to failures and lead recovery efforts.
  • Maintain compliance with industry standards and regulations.
  • Collaborate with Operations, Engineering, Facilities, and Vendors to integrate reliability practices.
  • Propose continuous-improvement initiatives and pilot emerging reliability technologies.

Skills

Analytical thinking
Communication
Project coordination
Team collaboration
Scripting

Education

Bachelor's degree in Engineering

Tools

Python
R
Monitoring platforms
IoT sensors

Job description

Khazna was founded in 2012 and has grown rapidly into becoming the leading and trusted wholesale Data Center provider in the Middle East and North Africa region. Through our Data Centers, we provide industry benchmark levels of power supply and cooling services to better serve the growing need for data center operations in the UAE and wider region.

We are seeking a Site Reliability Engineer to support the reliability engineering program across multiple data centers in our fleet. Reporting to the Reliability Manager, you will be responsible for monitoring system performance, driving preventative and predictive maintenance initiatives, leading root cause analysis efforts, and collaborating with cross-functional teams to minimize downtime and enhance infrastructure resilience.

Key Accountabilities:
  • Monitor real-time and historical performance metrics for critical power, cooling, and IT systems.
  • Analyse system data to identify trends, failure modes, and reliability risks.
  • Execute Root Cause Analyses (RCA) and Failure Mode & Effects Analyses (FMEA), then drive corrective and preventive actions.
  • Develop and maintain condition-based and predictive maintenance routines, leveraging IoT, data analytics, and machine learning tools.
  • Support preventive maintenance programs: schedule, document, and validate maintenance activities.
  • Assist in asset lifecycle planning, including upgrades, decommissioning, and end-of-life strategies.
  • Contribute to capacity runway assessments to forecast infrastructure needs.
  • Implement and enforce availability management plans, risk assessments, and mitigation strategies.
  • Ensure data collection and reporting processes for reliability KPIs (e.g., MTBF, MTTR, availability) are standardized and accurate.
  • Prepare reliability reports and dashboards; present findings and recommendations to site leadership.
  • Respond to and lead failure-response efforts during site incidents, ensuring rapid recovery and root-cause follow-through.
  • Maintain compliance with industry standards and regulations (Uptime Institute, ISO, ASHRAE).
  • Collaborate with Operations, Engineering, Facilities, and Vendors to integrate reliability best practices into day-to-day workflows.
  • Propose continuous-improvement initiatives and pilot emerging reliability technologies.
  • The job holder may be required to undertake additional duties, which may be reasonably expected and forms part of the function of the job.
Minimum Qualifications:
  • Bachelor’s degree in mechanical, Electrical, Reliability, or related Engineering discipline.
  • 3+ years of experience in reliability engineering, maintenance engineering, or a data center operations environment.
  • Hands‑on experience with RCA, FMEA, and predictive maintenance methodologies.
  • Proficiency with monitoring platforms, data‑analytics tools, and scripting (e.g., Python, R).
  • Familiarity with IoT sensors, machine-learning frameworks, and condition-based monitoring systems.
  • Knowledge of industry reliability standards and regulations (ISO, ASHRAE, Uptime Institute).
Job-Specific Skills (Generic / Technical):
  • Strong analytical and problem-solving skills, with acute attention to detail.
  • Effective communicator, able to present technical findings to diverse audiences.
  • Project coordination skills and the ability to manage multiple reliability initiatives.
  • Collaborative mindset, comfortable working in cross-functional teams.
  • Self-starter with a continuous-improvement attitude and commitment to resilience.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Reliability Engineer - Mechanical
Reliability Engineer - Mechanical

Tanqeeb • Dubai

On-site
AED 270,000 - 450,000
Data Center Reliability Engineer | Predictive Maintenance
Data Center Reliability Engineer | Predictive Maintenance

Khazna Data Centers • Dubai

On-site
AED 180,000 - 300,000
Data Center Site Reliability Engineer – Uptime & Resilience
Data Center Site Reliability Engineer – Uptime & Resilience

Tanqeeb • Dubai

On-site
AED 270,000 - 450,000
Site Energy Engineer
Site Energy Engineer

Khazna Data Centers • Abu Dhabi

On-site
AED 180,000 - 240,000
Manager - Design Mechanical
Manager - Design Mechanical

Khazna Data Centers • Dubai

On-site
AED 250,000 - 350,000
Senior Manager Implementation
Senior Manager Implementation

Khazna Data Centers • Abu Dhabi

On-site
AED 420,000 - 640,000
Manager – Mechanical Implementation
Manager – Mechanical Implementation

Khazna Data Centers • Abu Dhabi

On-site
AED 279,000 - 446,000
Senior Reliability Engineer
Senior Reliability Engineer

Star Services LLC • Abu Dhabi

On-site
AED 477,414 - 587,587
Quality & Reliability Manager
Quality & Reliability Manager

KBR, Inc. • Dubai

On-site
AED 300,000 - 400,000
Development Manager
Development Manager

Khazna Data Centers • Abu Dhabi

On-site
AED 350,000 - 550,000