Senior Platform Support Engineer (Remote)

Radian Arc

Kuala Lumpur

Hybrid

MYR 120,000 - 180,000

Full time

16 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid-friendly work model
International team environment
Career growth opportunities

Job summary

Radian Arc in Malaysia is seeking a Senior Cloud Support Engineer to provide advanced technical support for customers running GPU-based workloads in a hybrid environment. You will diagnose incidents across compute, networking, and storage, perform RCA, and collaborate with engineering and operations to improve platform reliability.

The role requires 5+ years in cloud support or similar, strong Linux and Kubernetes skills, and experience with GPU infrastructure and high-performance networking.

Qualifications

  • 5+ years of cloud support, infrastructure operations or systems administration.
  • Strong Linux systems administration skills and experience with Kubernetes.
  • Experience supporting GPU clusters and HPC environments.
  • Solid understanding of networking fundamentals for distributed workloads.
  • Proven incident response, RCA and post-mortem practices.
  • Ability to mentor others and work with cross-functional teams.

Responsibilities

  • Provide advanced support for customers operating GPU clusters and workloads.
  • Diagnose incidents across compute, networking, storage, and orchestration layers.
  • Perform root cause analysis and contribute to long-term remediation.
  • Escalate and coordinate with platform engineering and operations teams.
  • Improve runbooks, playbooks, and monitoring to reduce outages.
  • Mentor medior engineers and share troubleshooting expertise.

Skills

5+ years cloud support
Linux administration
Kubernetes
GPU infrastructure
Networking fundamentals
Incident response

Tools

Jira
Confluence
Zendesk
PagerDuty
Slack

Job description

About Radian Arc

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

About Radian Arc

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What Impact You Will Have

Mission: Provide advanced technical support for customers running workloads on the GPU cloud platform, ensuring reliable operation of edge- and large-scale GPU clusters and infrastructure services.

The Senior Cloud Support Engineer acts as a technical escalation point for complex incidents, helping diagnose and resolve issues across compute, networking, storage, and orchestration layers. This role works closely with engineering and operations teams to improve platform reliability, reduce incident frequency, and enhance the overall customer experience.

What You’ll Do

Customer Support & Incident Management

  • Provide advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure.
  • Diagnose and resolve complex customer issues affecting GPU clusters, compute nodes, networking, and storage.
  • Investigate incidents across multiple layers of the stack including firmware, drivers, operating systems, and platform services.
  • Perform root cause analysis (RCA) for major incidents and contribute to long-term remediation efforts.
  • Serve as a technical escalation point for complex or high-priority support cases.

Infrastructure Troubleshooting

  • Troubleshoot issues affecting:
    • GPU compute nodes
    • Kubernetes clusters
    • Networking infrastructure
    • Local NVMe, hyperconverged and distributed storage systems
  • Analyze logs, telemetry, and monitoring signals to identify underlying causes of platform instability.
  • Investigate issues related to GPU drivers, firmware, networking, and system performance.

Security Monitoring & Incident Triage

  • Monitor and investigate security alerts generated by the platform security stack.
  • Analyze and triage alerts generated by
    • Wazuh
    • TheHive
    • Cortex
  • Validate alerts, determine impact, and elevate potential security incidents to the security engineering team.
  • Assist in collecting system telemetry, logs, and forensic data required for incident investigations.
  • Improve alert runbooks and operational procedures to reduce false positives and improve response time.

GPU HPC Workload Support

  • Provide advanced support for large-scale GPU workloads running distributed training and inference jobs.
  • Diagnose failures affecting multi-GPU and multi-node workloads.
  • Investigate performance issues impacting distributed workloads, including:
    • GPU utilization
    • Communication latency
    • Storage bottlenecks
    • networking congestion
  • Support scheduling systems used for GPU workloads, including troubleshooting:
    • Job queue failures
    • Scheduling constraints
    • Cluster resource fragmentation

High-Performance Networking Troubleshooting

  • Diagnose issues affecting high-performance networking fabrics used by distributed workloads.
  • Support environments using:
    • RDMA
    • RoCE networking
  • Investigate performance issues affecting GPU-to-GPU communication and distributed training pipelines.

Data Center Coordination

  • Coordinate with customer’s data center technicians and infrastructure teams to perform remote diagnostics and hardware interventions when required.
  • Assist in validating on-premise installations and deployments of GPU infrastructure, ensuring hardware, networking, and platform components are correctly installed and operational.
  • Support hardware troubleshooting and identify faulty components across GPU nodes, networking equipment, and storage systems.
  • Coordinate and track hardware replacements and RMA processes with vendors and data center staff.
  • Validate hardware health after replacements, including GPU nodes, NICs, DPUs, storage devices, and power components.
  • Work closely with deployment and infrastructure teams to verify service readiness after installations, expansions, or hardware maintenance activities.

Operational Excellence

  • Participate in on-call 24/7 rotations to ensure production platform availability.
  • Respond to monitoring alerts and resolve operational incidents in accordance with defined SLAs.
  • Improve operational runbooks, troubleshooting guides, and support documentation.
  • Contribute to improving incident response processes and operational tooling.

Cross-Team Collaboration

  • Work closely with platform engineering, infrastructure engineering, and networking teams to resolve systemic issues.
  • Provide feedback to engineering teams on recurring operational problems affecting customers.
  • Help translate customer issues into actionable improvements for the platform.

Automation & Tooling

  • Develop automation scripts and tools to streamline support workflows.
  • Improve observability dashboards and alerts to enable faster issue detection and resolution.
  • Contribute to automation initiatives that reduce manual intervention in operational processes.

Knowledge Sharing & Mentorship

  • Mentor medior support engineers and share troubleshooting expertise across the team.
  • Lead the creation of knowledge base articles, troubleshooting guides, and operational documentation.
  • Contribute to training initiatives that improve the team’s technical capabilities.
Technical Stack
Operating Systems
  • Linux (Ubuntu)
GPU Infrastructure
  • NVIDIA GPU platforms
  • CUDA drivers
  • GPU monitoring tools (nvidia-smi)
Platform Infrastructure
  • Kubernetes
  • Container runtimes
  • KubeVirt
  • Distributed compute environments
Networking
  • NVIDIA Cumulus
  • OOB, north-south, and east-west fabric topologies
  • TCP/IP
  • VLAN, VXLAN, OVS/OVN
  • Routing fundamentals (BGP, VRFs)
  • DNS / DHCP
  • High-performance networking (RDMA/NVLink/NCCL)
Observability
  • Grafana
  • Zabbix
Security Monitoring
  • Wazuh
  • TheHive
  • Cortex
Automation
  • Python
  • Bash
  • Ansible, Terraform
Collaboration & Documentation
  • Jira
  • Confluence
  • Zendesk
  • PagerDuty
  • Slack
What You’ll Need
Core Experience
  • 5+ years of experience in cloud support, infrastructure operations, or systems administration.
  • Experience supporting large-scale infrastructure environments or GPU clusters.
Systems Expertise
  • Strong Linux systems administration skills.
  • Experience troubleshooting issues across compute, networking, and storage layers.
  • Familiarity with Kubernetes platforms and containerized workloads.
GPU Infrastructure
  • Experience working with GPU hardware platforms or HPC environments.
  • Familiarity with GPU monitoring tools and debugging GPU-related issues.
Networking
  • Solid understanding of networking fundamentals including L2/L3 concepts, routing, and load balancing.
  • Ability to diagnose connectivity issues affecting distributed workloads.
Operational Mindset
  • Strong troubleshooting and incident response skills.
  • Experience participating in on-call rotations and handling production incidents.
  • Ability to perform root cause analysis and drive operational improvements.
Communication & Collaboration
  • Excellent written and verbal communication skills.
  • Ability to explain complex technical concepts to both technical and non-technical stakeholders.
  • Proven ability to collaborate effectively with cross-functional engineering teams.

Location & work modality: Malaysia or comparable time zone

Start: August 2026

Type of Contract: Contractor

Average 40 hours per week, 9x5 business hour support with after hour on-call response/resolution for category 1 incidents

What We Offer
  • Attractive compensation package reflecting your expertise and experience.
  • A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
  • You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.

Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility

Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Platform Support Engineer – Remote-Ready
Senior GPU Platform Support Engineer – Remote-Ready

Radian Arc • Kuala Lumpur

Hybrid
MYR 120,000 - 180,000
Hybrid-friendly work model
International team environment
Career growth opportunities
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Senior Data Centre Operations Engineer
Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer Technologies Group • Cyberjaya

On-site
MYR 60,000 - 100,000
Data Centre GPU Infrastructure Engineer (Based KL)
Data Centre GPU Infrastructure Engineer (Based KL)

CloudEngine Digital Co., Ltd • Kuala Lumpur

On-site
MYR 67,000 - 112,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Kulai

On-site
MYR 60,000 - 120,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 90,000 - 150,000
Data Center Operations Engineer
Data Center Operations Engineer

Bitdeer (NASDAQ: BTDR) • Johor Bahru

On-site
MYR 100,000 - 145,000
Open workspaces
Training and mentoring
Welfare benefits
Senior Security Operations Engineer, AIDC
Senior Security Operations Engineer, AIDC

Bitdeer (NASDAQ: BTDR) • Penang

On-site
MYR 275,157 - 353,773
Attractive welfare benefits
Personal accountability and growth opportunities
Training and mentoring programs
System Engineer – Infrastructure (GPU - HPC Systems)
System Engineer – Infrastructure (GPU - HPC Systems)

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 120,000 - 200,000