Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Neuron Solutions Sdn. Bhd. in Malaysia seeks an experienced Shift Lead for Data Centre Systems Operations to lead on-shift GPU cluster and infrastructure activities.
You will manage incidents, ITSM governance and vendor coordination while ensuring service levels and stakeholder communications. The role requires a strong background in data centre operations, GPU hardware familiarity, and experience with Jira Service Management.
We are looking for an experienced Shift Lead – Data Centre Systems Operations to lead on-shift operations supporting large-scale GPU clusters and data centre infrastructure.
This role will be responsible for operational execution, incident coordination, ITSM governance, vendor coordination and customer-facing communication during the assigned shift.
Lead daily data centre and infrastructure operations during the assigned shift.
Oversee the stable operation of GPU clusters, networks, storage and supporting infrastructure.
Act as the primary escalation point for L1 engineers.
Coordinate P1/P2 and major incidents and ensure timely resolution.
Ensure incidents, tickets and escalations meet required SLA and ITSM standards.
Review shift handovers and ensure operational continuity.
Coordinate with hardware, network and data centre vendors.
Track vendor cases, SLA performance and hardware replacement activities.
Review operational dashboards and ensure data accuracy.
Ensure activities are properly documented and supported by evidence.
Prepare shift reports covering incidents, SLA risks, vendor cases and data centre activities.
Coach and guide L1 engineers on troubleshooting and operational processes.
Develop and maintain SOPs and operational runbooks.
Bachelor's degree in Computer Science, IT, Electrical Engineering or a related field, or equivalent practical experience.
5+ years of experience in IT infrastructure or data centre operations.
Previous experience as a Shift Lead, Senior NOC Engineer, Operations Lead or similar role.
Strong ITSM and incident management experience.
Experience with Jira Service Management or similar ticketing platforms.
Experience managing P1/P2 incidents and SLA-driven environments.
Familiarity with NVIDIA GPU hardware and AI/HPC environments is an advantage.
Experience with monitoring tools such as Prometheus, Grafana or Zabbix.