An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Achieva Technology Sdn Bhd is seeking a Data Centre Project Manager and On Site Engineering Lead to run the on‑site engineering team at our Malaysia data centre. You will install servers, diagnose hardware faults and manage 24/7 incident response to meet SLAs.
You will coordinate with data centre providers, plan shifts, oversee installation of GPU servers and infrastructure, and drive project milestones from site prep to handover. Strong documentation and multilingual communication are required.
Company: Achieva Cloud Services | Location: Malaysia, based on site at the data centre | Employment type: Full time
We are looking for a hands-on leader to manage the deployment and ongoing operation of our AI computing infrastructure at the data centre. You will lead the on-site engineering team, organise 24/7 coverage and incident response to meet customer service level agreements (SLAs). You must be able to install servers, diagnose hardware faults and replace failed components yourself when needed. You will also manage the data centre provider’s delivery schedule and verify that its technical specifications meet Achieva Cloud’s requirements.
Lead, coach and supervise the on-site engineers, with clear ownership of daily tasks, maintenance activities and incident response.
Plan shift rosters and handovers to maintain 24/7 coverage, including suitable cover for leave, training and absences.
Own operational delivery against agreed customer SLAs, including monitoring, response, escalation, resolution coordination and service reporting.
Set incident priorities, coordinate the response across shifts, and elevate major incidents promptly to management, customers and relevant vendors.
Review team performance, incident trends and recurring issues; implement corrective actions and maintain runbooks and training records.
Plan and manage the installation of GPU servers, racks, networking and related infrastructure, from site preparation through testing and handover.
Coordinate delivery schedules, site access, installation teams, vendors and customer requirements.
Track project milestones, risks, costs and outstanding issues; provide regular progress updates to management.
Verify rack space, allocated power, connectivity and cooling readiness with the data centre operator before equipment installation.
Maintain installation records, asset registers, rack layouts, network diagrams and handover documents.
Conduct routine checks on server and network equipment, rack conditions, alerts and capacity utilisation.
Personally install and commission servers, network equipment and cabling when required, and supervise the team’s installation work.
Diagnose hardware faults using system alerts, logs and management interfaces; replace failed field-replaceable components and verify that service is restored.
Perform hot‑swap replacement of supported parts, such as drives, power supplies and fans, according to the manufacturer’s procedures and approved change controls.
Coordinate complex repairs, warranty claims, spare parts and firmware updates with equipment vendors and internal technical teams.
Direct the team during incidents and coordinate recovery, customer updates and root‑cause analysis.
Manage planned changes using documented approvals, maintenance windows and rollback procedures.
Monitor available rack space, power allocation and deployment capacity to support future expansion.
Act as Achieva Cloud’s primary operational contact for the data centre provider, equipment suppliers, contractors and service providers.
Review the provider’s proposed design and specifications against agreed requirements for rack capacity and dimensions, power density and redundancy, cooling, connectivity, physical security and site access.
Track the provider’s buildout, commissioning and handover milestones; inspect progress, document gaps and follow through on corrective actions before accepting the space for deployment.
Coordinate with the provider on available capacity, power and cooling performance, maintenance schedules and facility incidents throughout operations.
Monitor the provider’s performance against its contractual commitments and expedite delays, specification deviations and service issues to management.
Oversee vendor work on site and confirm that installations meet agreed technical, safety and security requirements.
Support customer visits, equipment acceptance and technical audits when required.
Maintain incident, change, maintenance and access records.
Track SLA performance, follow up on recurring faults and corrective actions, and report any service risks or breaches.
Prepare concise operational reports covering project status, infrastructure health, incidents and capacity.
Participate in the management escalation rota for major incidents outside normal working hours.
Degree in IT, computer engineering, electrical engineering, telecommunications or a related technical field.
At least five years of relevant experience, including two years in a data centre or other critical infrastructure environment.
Proven hands‑on ability to install, configure and troubleshoot servers, network equipment and structured cabling.
Experience diagnosing hardware failures and replacing server components, including safe hot‑swap procedures for supported parts.
Familiarity with data centre power, cooling, physical security and change‑control processes.
Ability to manage multiple vendors and projects while remaining comfortable troubleshooting issues on site.
Ability to review data centre technical specifications, track provider milestones and identify gaps against agreed requirements.
Experience supervising technical staff, organising shift coverage and managing incident escalations in a 24/7 service environment.
Clear communication, strong documentation habits and sound judgment during incidents.
Fluent in spoken and written English and Mandarin to communicate with Mandarin speaking customers, vendors and technical teams.
Willingness to work on site in Malaysia and attend planned maintenance or critical incidents outside normal hours.
The provider meets agreed specifications and milestones, the engineering team sustains 24/7 SLA coverage, and hardware faults are resolved promptly and safely.
Established in year 1996, Achieva Technology Sdn Bhd was incorporated as a subsidiary of Achieva Technologies Pte Ltd in Singapore under Achieva Limited* umbrella.
ATSB has been a distributor bridging domestic needs to that of the foreign manufacturers. The core business is the distribution of Computer and Computer Peripheral products.
*Achieva Limited was Public Listed in Singapore since Year 2000. Geographical coverage of Achieva Group of Companies includes Singapore, Malaysia, Indonesia, Philippines, Vietnam and Australia. Market capitalization is S$56 million as at 7th August 2009 *
To help fast track investigation, please include here any other relevant details that prompted you to report this job ad as fraudulent / misleading / discriminatory / salary below minimum wage.
Researching careers? Find all the information and tips you need on career advice.