Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc.

Kuala Lumpur

On-site

MYR 180,000 - 280,000

Full time

4 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Risewave Consulting, Inc. in Kuala Lumpur is seeking an experienced Platform Engineer to lead end-to-end GPU compute architecture, HPC and AI infrastructure initiatives for enterprise workloads.

You will drive Kubernetes-based GPU scheduling, high-speed interconnects, and scalable cloud services while collaborating with product and sales teams to define roadmaps. The role demands hands-on GPU cluster experience, Linux/Kubernetes proficiency, and the ability to mentor a growing team as the

Qualifications

  • Bachelor’s degree or above in Computer Science, Electronics, or related field.
  • 5+ years in Cloud Computing, HPC, AI Infrastructure, or Data Centre Engineering.
  • Hands-on 2+ years with GPU clusters, AI cloud platforms, or large-scale compute environments.
  • Strong Linux, Docker, Kubernetes, Helm, and cloud-native ecosystems knowledge.
  • Experience with Kubernetes GPU scheduling or Slurm cluster management.

Responsibilities

  • Drive end-to-end architectural design and reviews for GPU servers and AI compute clusters.
  • Assess hardware, interconnects, storage, and network topologies for performance.
  • Lead platform operations, SLA management, and multi-tenant resource governance.
  • Deploy and maintain GPU resource scheduling/orchestration platforms (Kubernetes/Slurm).
  • Collaborate with product and sales to translate requirements into technical roadmaps.
  • Support LLM training, fine-tuning, and inference deployment for enterprise workloads.

Skills

Cloud computing
HPC
AI infrastructure
Linux
Kubernetes
Docker
Vendor management
Bilingual (English/Mandarin)
Communication
Team leadership

Education

Bachelor's degree in CS/EE/related

Tools

Kubernetes scheduling
Slurm
DCGM
Prometheus
Grafana

Job description

  • Drive end-to-end technical architecture design and review for GPU servers and AI compute clusters.
  • Participate in the evaluation and selection of mainstream high-performance GPU hardware platforms and server architectures.
  • Assess server topologies, CPUs, GPUs, system memory, high-speed local NVMe storage, and network interfaces.
  • Design and review high-speed networking fabrics, including InfiniBand, RoCE, and high-speed Ethernet topologies.
  • Evaluate NVLink, NVSwitch, RDMA, and multi-node/multi-GPU interconnect architectures.
  • Work with data centre teams to define power density, liquid/air cooling, PDU distribution, network topology, and cabling specifications.
  • Establish standard operating procedures for cluster deployment, benchmark testing, acceptance, and expansion.
  • Steer the technical strategy and buildout of GPU resource pools, compute clusters, and cloud platform services.
  • Deploy and maintain GPU resource scheduling and orchestration platforms built on Kubernetes and/or Slurm.
  • Manage GPU drivers, parallel acceleration libraries, container runtimes, GPU Operators, and underlying software stacks.
  • Architect multi-tier product offerings spanning Bare Metal, Virtual Machines, Containerized Instances, and On-Demand GPU slots.
  • Implement full-card allocation, MIG slicing, GPU sharing, and multi-tenant resource isolation.
  • Advance the development of user access control, quota management, automated provisioning, billing integration, and API services.
  • Collaborate closely with product and business teams to bring standardized GPU Cloud products to market.
  • Support enterprise customers with LLM training, fine-tuning, inference deployment, and performance optimization.
  • Analyze customer workload specifications, including model parameter scales, dataset size, concurrency, throughput, and latency SLAs.
  • Recommend optimal GPU models, cluster node counts, network fabrics, and storage configurations tailored to client workloads.
  • Support mainstream AI frameworks (PyTorch, TensorFlow) and distributed training frameworks.
  • Support high-performance inference frameworks and microservices such as vLLM, Triton, TensorRT-LLM, and AI microservice architectures.
  • Lead proofs-of-concept (PoCs), performance benchmarking, and technical acceptance testing.
  • Platform Operations & SLA Management
  • Build comprehensive monitoring and alerting systems across GPUs, compute nodes, network switches, and storage (e.g., DCGM, Prometheus, Grafana).
  • Track GPU utilization, VRAM usage, power consumption, thermal profiles, ECC hardware errors, and idle resource rates.
  • Implement metering logic based on GPU-hours, instance-hours, or token generation metrics.
  • Define Service Level Agreements (SLAs), incident severity levels, escalation pathways, and disaster recovery plans.
  • Continuously optimize platform uptime, compute efficiency, and commercial output per GPU node.
  • Partner with sales and business development teams during customer technical discovery calls and solution proposals.
  • Translate client AI requirements into technical architecture blueprints, Bills of Materials (BOMs), RFPs/RFQs, and execution roadmaps.
  • Deliver technical presentations, platform demonstrations, PoCs, and technical onboarding for strategic clients.
  • Define project technical boundaries, deliverables, SLAs, and customer acceptance criteria.
  • Manage technical delivery with server OEMs, data centre operators, network/storage vendors, and system integrators.
  • Oversee hardware delivery, rack mounting, cabling, environment initialization, cluster commissioning, and acceptance.
  • Manage project timelines, mitigate technical risks, and drive issue resolution.
  • Standardize technical documentation, operational runbooks, and delivery guidelines.
  • Build and lead a team of infrastructure and platform engineers as the business scales.
Job Requirements
  • Bachelor’s degree or above in Computer Science, Electronic Engineering, Telecommunications, Software Engineering, or a related discipline.
  • A minimum of 5 years of experience in Cloud Computing, High-Performance Computing (HPC), AI Infrastructure, or Data Centre Engineering.
  • At least 2 years of hands‑on experience with GPU clusters, AI cloud platforms, or large‑scale compute environments.
  • Strong proficiency in Linux, Docker, Kubernetes, Helm, and cloud‑native container ecosystems.
  • In‑depth knowledge of GPU hardware architectures, parallel compute libraries, driver stacks, and AI server architectures.
  • Hands‑on experience with Kubernetes GPU scheduling or Slurm cluster management.
  • Solid understanding of high‑speed interconnect technologies, including InfiniBand, RoCE, RDMA, and NVLink.
  • Practical understanding of AI workload execution spanning LLM training, fine‑tuning, and inference serving.
  • Strong client‑facing communication, cross‑functional collaboration, and vendor management skills.
  • Fluent in both English and Mandarin, with the ability to conduct technical meetings and produce technical documentation in both languages.
  • Based in Malaysia with willingness to travel within the Southeast Asia region in future.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Engineer – Infrastructure (AI & HPC Systems)
System Engineer – Infrastructure (AI & HPC Systems)

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Senior Data Centre Operations Engineer
Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Kulai

On-site
MYR 60,000 - 120,000
Senior AI Data Centre Network Engineer
Senior AI Data Centre Network Engineer

YTL AI Cloud • Kulai

On-site
MYR 180,000 - 280,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 90,000 - 150,000
Senior AI Network - Security Engineer
Senior AI Network - Security Engineer

Techstreet Malaysia • Johor Bahru

On-site
MYR 180,000 - 280,000
Senior AI Network & Security Engineer (Johor Bahru)
Senior AI Network & Security Engineer (Johor Bahru)

Techstreet • Johor Bahru

On-site
MYR 180,000 - 300,000
AI Infra Engineer: GPU Cloud & ML Platform
AI Infra Engineer: GPU Cloud & ML Platform

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000
AI Engineer SNS Network Right Choice with the Right People
AI Engineer SNS Network Right Choice with the Right People

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000
AI Infrastructure & Orchestration Lead
AI Infrastructure & Orchestration Lead

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000