Get more replies from employers
Send a job-specific resume in minutes.
Neuron Solutions Sdn. Bhd.
in Johor Bahru, Johor is seeking a hands-on System Engineer - Infrastructure to design, deploy and optimise large-scale GPU and CPU compute environments powering AI workloads and high‑performance computing. You will manage server provisioning, OS deployment, hardware diagnostics and lifecycle management, while collaborating with networking and DevOps teams to ensure reliability, performance and scalable operations.
Neuron Solutions Sdn. Bhd. - Johor Bahru, Johor
We are looking for a hands‑on System Engineer - Infrastructure to support and optimise large-scale GPU and CPU infrastructure powering AI workloads, large model training and high-performance computing environments.
You will be responsible for deploying, maintaining and troubleshooting GPU/CPU servers while ensuring infrastructure reliability, performance and operational readiness.
Deploy, configure and maintain GPU and CPU servers across large-scale compute clusters.
Optimise BIOS, firmware and operating system configurations for AI and HPC workloads.
Perform server health monitoring, hardware diagnostics, firmware upgrades and lifecycle management.
Support cluster deployments across multiple racks and coordinate with hardware vendors and system integrators.
Manage server provisioning, OS deployment, GPU/NIC driver installation and system hardening.
Conduct system validation, burn‑in testing and workload benchmarking.
Monitor system health, investigate failures and perform root cause analysis.
Work closely with networking, storage and DevOps teams to ensure end‑to‑end infrastructure performance.
Develop scripts and automation to improve infrastructure deployment, monitoring and remediation.
Maintain technical documentation including system configurations, rack layouts, cabling and operational procedures.
Provide L2/L3 support and participate in on‑call activities.
Bachelor's degree in Computer Science, Electrical Engineering or a related technical field.
At least 3 years of hands‑on experience managing server infrastructure in HPC, AI, GPU cluster or data centre environments.
Strong experience with Linux systems, system tuning and performance optimisation.
Hands‑on experience with GPU/CPU servers and bare‑metal infrastructure.
Knowledge of server provisioning technologies such as IPMI, PXE, Redfish or BMC.
Familiarity with monitoring tools such as Prometheus and Grafana.
Basic knowledge of Kubernetes or containerised environments.
Experience with server hardware, GPU platforms and infrastructure troubleshooting.