Senior Infrastructure Engineer

MTAI Sdn. Bhd.

Kuala Lumpur

On-site

MYR 180,000 - 300,000

Full time

3 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

MTAI Sdn. Bhd. seeks an experienced Infrastructure/Datacenter Engineer to own remote management of NVIDIA GPU servers and HPC fabric. You will coordinate hardware vendors, manage lifecycle processes and help scale the platform with HPC interconnects and orchestration tooling.

Ideal candidates have 8+ years in a similar role, hands-on GPU HPC ops, and strong Linux expertise. The role involves startup-like ownership in a small, fast-moving team in Kuala Lumpur.

Qualifications

  • 8+ years in infrastructure, systems or datacenter engineering.
  • Hands-on remote datacenter server management experience.
  • Experience managing hardware/OEM vendor relationships.
  • Fluency in HPC environments and InfiniBand interconnects.

Responsibilities

  • Own remote management, monitoring and troubleshooting of datacenter servers, primarily NVIDIA GPU servers.
  • Manage datacenter, hardware and OEM vendor relationships including procurement and support escalation.
  • Support HPC infrastructure, including high-speed interconnects (InfiniBand) at fabric level.
  • Deploy, configure and operate Kubernetes and Slurm for GPU workloads.
  • Define standards and processes for hardware provisioning and server lifecycle management.
  • Diagnose hardware, firmware and low-level infra issues with vendor coordination.
  • Collaborate with Head of Sovereign AI Infra and infra team to ensure integrated stack.
  • Contribute to capacity planning and scaling as platform grows.

Skills

Datacenter management
Vendor management
Kubernetes
Slurm
GPU compute infra
Linux
InfiniBand HPC networking
Firmware management
Capacity planning
Automation

Job description

  • Own the remote management, monitoring and troubleshooting of our datacenter servers, primarily NVIDIA GPU servers (DGX/HGX-class or similar).
  • Manage relationships with datacenter, hardware and OEM vendors, including procurement, support escalation and lifecycle management.
  • Support and maintain high-performance computing (HPC) infrastructure, including high-speed interconnects (InfiniBand) at the fabric level.
  • Support the deployment, configuration and day-to-day operation of orchestration platforms — Kubernetes and Slurm — for GPU workloads.
  • Define and establish standards and processes for hardware provisioning, server lifecycle management, and datacenter operations.
  • Diagnose and resolve hardware, firmware and low-level infrastructure issues, coordinating with vendors where necessary.
  • Work closely with the Head of Sovereign AI Infra and the broader infrastructure team to ensure hardware, orchestration and platform layers work together reliably.
  • Contribute to capacity planning and infrastructure scaling decisions as the platform grows.
  • 8+ years of experience in an infrastructure, systems or datacenter engineering role.
  • Hands-on experience with remote management of datacenter servers, including hardware troubleshooting, firmware/BIOS management and vendor coordination.
  • Experience managing hardware/OEM vendor relationships.
  • Fluency in high-performance computing (HPC) environments and high-speed interconnects (InfiniBand) — hands-on experience is ideal, but strong working fluency is acceptable.
  • Experience with orchestration platforms for compute clusters, particularly Kubernetes and Slurm.
  • Demonstrated ability to define and introduce standards and processes from scratch, ideally in a startup or early-stage environment.
  • Comfortable working in a small, fast-moving team with significant ownership.

Highly Regarded

  • Direct hands-on experience with NVIDIA GPU server platforms (DGX, HGX or similar).
  • Experience supporting AI/ML training or inferencing workloads at the infrastructure level.
  • Familiarity with Linux and the open-source infrastructure ecosystem.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

iSoftStone • Kuala Lumpur

On-site
MYR 60,000 - 100,000
Senior AI Network & Security Engineer
Senior AI Network & Security Engineer

techstreet • Johor

On-site
MYR 180,000 - 300,000
Health Insurance
Performance Bonus
Dental Coverage
AI Network & Security Engineer - Data Centre
AI Network & Security Engineer - Data Centre

Neuron Solutions Sdn. Bhd. • Johor

On-site
MYR 120,000 - 180,000
Data Centre Infrastructure Engineer
Data Centre Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000
Shift Lead – DC Systems Operations Engineer
Shift Lead – DC Systems Operations Engineer

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 120,000 - 190,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer Technologies Group • Cyberjaya

On-site
MYR 60,000 - 100,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer • Cyberjaya

On-site
MYR 67,000 - 100,000
System Engineer – Infrastructure (AI & HPC Systems)
System Engineer – Infrastructure (AI & HPC Systems)

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Senior AI Network & Security Engineer (Johor Bahru)
Senior AI Network & Security Engineer (Johor Bahru)

Techstreet • Johor Bahru

On-site
MYR 180,000 - 300,000