System Engineer – Infrastructure (AI & HPC Systems)

Neuron Solutions Sdn. Bhd.

Johor Bahru

On-site

MYR 90,000 - 150,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Monetary compensation

Job summary

Neuron Solutions Sdn. Bhd.

in Johor Bahru, Johor is seeking a hands-on System Engineer - Infrastructure to design, deploy and optimise large-scale GPU and CPU compute environments powering AI workloads and high‑performance computing. You will manage server provisioning, OS deployment, hardware diagnostics and lifecycle management, while collaborating with networking and DevOps teams to ensure reliability, performance and scalable operations.

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering or a related technical field.
  • At least 3 years hands-on experience managing server infrastructure in HPC, AI, GPU cluster or data centre environments.
  • Strong experience with Linux systems, system tuning and performance optimisation.
  • Hands-on experience with GPU/CPU servers and bare-metal infrastructure.
  • Knowledge of server provisioning technologies such as IPMI, PXE, Redfish or BMC.
  • Familiarity with monitoring tools such as Prometheus and Grafana.
  • Basic knowledge of Kubernetes or containerised environments.
  • Experience with server hardware, GPU platforms and infrastructure troubleshooting.

Responsibilities

  • Deploy, configure and maintain GPU and CPU servers across large-scale compute clusters.
  • Optimise BIOS, firmware and operating system configurations for AI and HPC workloads.
  • Perform server health monitoring, hardware diagnostics, firmware upgrades and lifecycle management.
  • Support cluster deployments across multiple racks and coordinate with hardware vendors and system integrators.
  • Manage server provisioning, OS deployment, GPU/NIC driver installation and system hardening.
  • Conduct system validation, burn‑in testing and workload benchmarking.
  • Monitor system health, investigate failures and perform root cause analysis.
  • Work closely with networking, storage and DevOps teams to ensure end‑to‑end infrastructure performance.
  • Develop scripts and automation to improve infrastructure deployment, monitoring and remediation.
  • Maintain technical documentation including system configurations, rack layouts, cabling and operational procedures.
  • Provide L2/L3 support and participate in on‑call activities.

Skills

Linux systems
Server infrastructure
HPC/AI compute
System tuning
On-call support

Education

Bachelor's degree in Computer Science, Electrical Engineering or related technical field

Tools

IPMI
PXE
Redfish
BMC
Prometheus
Grafana
Kubernetes basics

Job description

Neuron Solutions Sdn. Bhd. - Johor Bahru, Johor

We are looking for a hands‑on System Engineer - Infrastructure to support and optimise large-scale GPU and CPU infrastructure powering AI workloads, large model training and high-performance computing environments.

You will be responsible for deploying, maintaining and troubleshooting GPU/CPU servers while ensuring infrastructure reliability, performance and operational readiness.

Key Responsibilities
  • Deploy, configure and maintain GPU and CPU servers across large-scale compute clusters.

  • Optimise BIOS, firmware and operating system configurations for AI and HPC workloads.

  • Perform server health monitoring, hardware diagnostics, firmware upgrades and lifecycle management.

  • Support cluster deployments across multiple racks and coordinate with hardware vendors and system integrators.

  • Manage server provisioning, OS deployment, GPU/NIC driver installation and system hardening.

  • Conduct system validation, burn‑in testing and workload benchmarking.

  • Monitor system health, investigate failures and perform root cause analysis.

  • Work closely with networking, storage and DevOps teams to ensure end‑to‑end infrastructure performance.

  • Develop scripts and automation to improve infrastructure deployment, monitoring and remediation.

  • Maintain technical documentation including system configurations, rack layouts, cabling and operational procedures.

  • Provide L2/L3 support and participate in on‑call activities.

Requirements
  • Bachelor's degree in Computer Science, Electrical Engineering or a related technical field.

  • At least 3 years of hands‑on experience managing server infrastructure in HPC, AI, GPU cluster or data centre environments.

  • Strong experience with Linux systems, system tuning and performance optimisation.

  • Hands‑on experience with GPU/CPU servers and bare‑metal infrastructure.

  • Knowledge of server provisioning technologies such as IPMI, PXE, Redfish or BMC.

  • Familiarity with monitoring tools such as Prometheus and Grafana.

  • Basic knowledge of Kubernetes or containerised environments.

  • Experience with server hardware, GPU platforms and infrastructure troubleshooting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Systems Engineer for AI & HPC Clusters
Infrastructure Systems Engineer for AI & HPC Clusters

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Senior Data Centre Operations Engineer
Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Kulai

On-site
MYR 60,000 - 120,000
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 90,000 - 150,000
Senior Data Center Ops Engineer — GPU & AI Infra
Senior Data Center Ops Engineer — GPU & AI Infra

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
Senior AI Data Centre Network Engineer
Senior AI Data Centre Network Engineer

YTL AI Cloud • Kulai

On-site
MYR 180,000 - 280,000
Senior AI Network & Security Engineer (Johor Bahru)
Senior AI Network & Security Engineer (Johor Bahru)

Techstreet • Johor Bahru

On-site
MYR 180,000 - 300,000
Senior AI Network - Security Engineer
Senior AI Network - Security Engineer

Techstreet Malaysia • Johor Bahru

On-site
MYR 180,000 - 280,000
AI/HPC Data Center Infrastructure Engineer
AI/HPC Data Center Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000