GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.
Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.
From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.
One cloud for compute, inference, and agents.
Role Overview
We are seeking a highly skilled and experienced DC Facility Lead Engineer (MEP - Mechanical, Electrical, and Plumbing) to join the GMI Global Infrastructure Team to oversee the operation and maintenance of the mechanical, electrical and plumbing (MEP) infrastructure for GMI global data centers. The ideal candidate will possess a strong technical background, critical thinking skills, and a commitment to safety protocols.
As the DC Facility Lead Engineer, you will be the backbone of our site uptime and operational excellence. You will be responsible for the end-to-end maintenance, optimization, and reliability of critical infrastructure, with a heavy emphasis on advanced cooling technologies (like CDUs) and robust power distribution (Busways). You will own the creation of operational documentation and lead routine maintenance planning to ensure 100% continuous availability for our high-density AI clusters.
Responsibilities
- Operations & Maintenance
- Oversee daily operations, routine testing, and preventative/predictive maintenance planning for all critical MEP infrastructure.
- Direct responsibility for specialized high-density AI data center equipment, including Busway systems, Cooling Distribution Units (CDUs) for liquid cooling, and Building Management Systems (BMS/DCIM).
- Act as the primary technical point of contact for incident management, troubleshooting complex MEP anomalies, and conducting Root Cause Analysis (RCA).
- Manage external vendors and contractors during routine servicing, ensuring all work meets rigorous safety and operational standards.
- Documentation & Process
- Author, review, and continuously optimize critical site documentation to ensure zero-downtime execution:
- Standard Operating Procedures (SOPs): Daily routines, equipment switching, and system configurations.
- Emergency Operating Procedures (EOPs): High-stress disaster recovery, power failures, and leak response protocols.
- Method of Procedures (MOPs): Step-by-step risk mitigation plans for live maintenance on critical systems.
- Establish and audit site safety standards, ensuring full compliance with local Taiwan regulations and global data center best practices.
- Monitoring & Capacity Management
- Monitor BMS/DCIM dashboards to track power usage effectiveness (PUE), liquid cooling flow rates, and thermal thresholds.
- Collaborate with the global infrastructure team to plan future capacity upgrades and efficiency optimizations.
Qualifications
- Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Facility Management, or a related technical field
- Minimum of 5 years’ experience in facility engineering, preferably within data centers MEP operations or critical facilities.
- Hands-on experience with high-density power distribution (Track Busways, RPPs, UPS systems).
- Strong working knowledge of liquid cooling loops, chillers, pumps, and Cooling Distribution Units (CDUs).
- Proficient in operating, configuring, and extracting data from BMS (Building Management Systems) or SCADA networks.
- Proven track record of independently writing and implementing high-quality MOPs, SOPs, and EOPs
- Excellent problem-solving skills and the ability to work under pressure.
- Strong communication skills to collaborate effectively with team members and stakeholders.
- Experience with energy management systems and sustainability practices.
- Familiarity with industry regulations.
- Ability to read and interpret technical documents and blueprints.
- Certified Data Center Professional (CDCP), Certified Data Center Specialist (CDCS), or equivalent industry certifications.
- Valid local professional licenses (e.g., Grade A/B Electrical Technician in Taiwan).
- Experience with high-density GPU/AI hardware clusters is preferred.
- Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.