Senior Data Center Infrastructure & Systems Engineer

Genesis Networks Pte Ltd

Iskandar Puteri

On-site

MYR 180,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Genesis Networks Pte Ltd seeks a Senior Data Center Infrastructure & Systems Engineer to lead deployment, network integration, and diagnostics for high-density GPU AI clusters in Malaysia. You will provision hardware, configure network fabrics, and validate performance across multi-node systems.

The role emphasizes Linux administration, PXE provisioning, GPU diagnostics, and SLAs, with responsibilities spanning burn-in, monitoring, incident response, and asset tracking in a fast-paced

Qualifications

  • Bachelor's Degree in Computer Engineering, Computer Science, Network Engineering, Systems Administration, or equivalent.
  • 3–5+ years of hands-on HPC, hyperscale data centers, or AI cluster infrastructure experience.
  • Experience with high-density GPU server hardware and liquid-cooled rack systems.
  • Strong Linux administration, PXE provisioning, and scripting skills.
  • Familiarity with GPU diagnostic tools (DCGM, nvidia-smi) and IPMI/Redfish APIs.
  • Familiarity with ticketing systems and asset tracking workflows.
  • Ability to adapt between fast-paced deployment and SLA-driven operations.

Responsibilities

  • Lead technical deployment, network integration, and diagnostic testing of GPU AI cluster infrastructure.
  • Configure OOB management networks, bare-metal OS provisioning, kernel tuning, and CUDA/driver installations.
  • Deploy and validate InfiniBand/RoCE fabrics, including switch config, BGP, and VXLAN overlays.
  • Perform multi-tier hardware validation: GPU diagnostics, memory bandwidth, power stress, multi-node benchmarks.
  • Conduct burn-in tests, monitor thermal thresholds, and support performance sign-off.
  • Handle SLA-driven work orders including hardware troubleshooting, swaps, and rack maintenance.
  • Monitor cluster health telemetry and proactively identify hardware issues.

Skills

Linux systems administration
Scripting (Bash/Python)
GPU diagnostics
Network engineering basics
Troubleshooting

Education

Bachelor's Degree in Computer Engineering, Computer Science, Network Engineering, Systems Administration, or equivalent

Tools

InfiniBand
RoCEv2
BGP
VXLAN
IPMI/Redfish
NVIDIA DCGM
nvidia-smi

Job description

We are looking for a Senior Data Center Infrastructure & Systems Engineer to lead the technical deployment, network integration, and diagnostic testing of high-density GPU AI cluster infrastructure. The role covers hardware provisioning, network fabric configuration, and rigorous performance validation, transitioning into ongoing monitoring, troubleshooting, and maintenance of the cluster fleet.

Key Responsibilities
  • Configure Out-of-Band (OOB) management networks and execute bare-metal OS provisioning, kernel tuning, and CUDA/driver installations across server nodes.
  • Deploy and validate high-speed network fabrics (InfiniBand, RoCEv2 Ethernet), including switch configuration, BGP routing, and VXLAN overlays.
  • Perform multi-tier hardware validation, including GPU diagnostics, memory bandwidth benchmarks, power stress tests, and multi-node performance benchmarking.
  • Conduct burn-in testing, monitor thermal thresholds, and support functional performance acceptance sign-off.
  • Execute SLA-driven work orders including hardware troubleshooting, component swaps, and rack-level maintenance.
  • Monitor cluster health telemetry (power, temperature, ECC errors, network performance) to proactively identify hardware issues.
  • Respond to infrastructure incidents, isolate faulty hardware, and support non-disruptive maintenance and firmware updates.
  • Maintain accurate asset inventory records and comply with physical security and data confidentiality policies.
Qualifications & Experience
  • Bachelor's Degree in Computer Engineering, Computer Science, Network Engineering, Systems Administration, or equivalent practical experience.
  • 3-5+ years of hands-on experience in high-performance computing (HPC), hyperscale data centers, or AI cluster infrastructure.
  • Proven experience with high-density GPU server hardware and liquid-cooled rack systems.
  • Strong Linux systems administration skills (Ubuntu/RHEL), PXE provisioning, and scripting (Bash/Python).
  • Expertise in high-speed network fabrics: InfiniBand, RoCEv2, BGP, VXLAN, IPAM.
  • Experience with GPU diagnostic tools (NVIDIA DCGM, nvidia-smi, Fabric Manager) and IPMI/Redfish APIs.
  • Familiarity with ticketing systems and asset tracking workflows.
  • Strong diagnostic and troubleshooting skills for complex hardware/network issues.
  • Strict adherence to safety and security protocols.
  • Ability to adapt between fast-paced deployment work and structured, SLA-driven operations.
Work Location

You will be based in Johor Bahru, within the Iskandar Puteri area.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Centre Operations Engineer
Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
System Engineer – Infrastructure (AI & HPC Systems)
System Engineer – Infrastructure (AI & HPC Systems)

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Data Center Project Manager (Deployment) / Operations Manager
Data Center Project Manager (Deployment) / Operations Manager

Genesis Networks Pte Ltd • Iskandar Puteri

On-site
MYR 180,000 - 260,000
Senior Data Center Ops Engineer — GPU & AI Infra
Senior Data Center Ops Engineer — GPU & AI Infra

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 90,000 - 150,000
System Engineer
System Engineer

Raydian Cloud Sdn Bhd • Kuala Lumpur

On-site
MYR 18,000 - 30,000
Senior AI Network & Security Engineer
Senior AI Network & Security Engineer

techstreet • Johor

On-site
MYR 180,000 - 300,000
Health Insurance
Performance Bonus
Dental Coverage
Senior GPU HPC Infra Engineer | Kubernetes & Slurm
Senior GPU HPC Infra Engineer | Kubernetes & Slurm

MTAI Sdn. Bhd. • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior AI Network & Security Engineer
Senior AI Network & Security Engineer

Techstreet Malaysia • Johor

On-site
MYR 150,000 - 230,000
Infrastructure Systems Engineer for AI & HPC Clusters
Infrastructure Systems Engineer for AI & HPC Clusters

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation