Technical Support Engineer (GPU/HPC)

Kerry Consulting

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Kerry Consulting in Singapore is seeking a Senior Technical Support Engineer for GPU and cloud infrastructure to resolve complex customer and production issues across GPU compute environments. You will collaborate with Engineering and SRE teams to strengthen platform reliability and own escalations.

You will lead root‑cause analysis, drive permanent solutions, and contribute to monitoring, runbooks, and incident response while participating in on‑call rotations.

Qualifications

  • 6+ years in cloud infra, escalation engineering, production ops, or SRE-adjacent roles.
  • Hands-on Linux, networking and cloud infra skills.
  • Experience with GPU infra, CUDA, HPC, or large-scale compute environments.

Responsibilities

  • Own complex escalations and platform incidents end-to-end.
  • Lead root-cause analysis and implement permanent solutions.
  • Improve monitoring, runbooks and engineering documentation.

Skills

Linux
Networking
Cloud infrastructure
GPU/CUDA

Tools

CUDA
GPU compute platforms

Job description

Overview

I'm partnering a rapidly expanding AI infrastructure and cloud computing company to hire a Senior Technical Support Engineer, GPU & Cloud Infrastructure in Singapore. This is a senior technical role focused on resolving complex customer and production issues across GPU compute environments while working closely with Engineering and SRE teams to strengthen platform reliability.

Responsibilities

You will take end-to-end ownership of complex escalations and platform incidents, troubleshooting across GPU compute, Linux, networking/SDN, storage, CUDA and drivers, control plane, and billing systems. Working closely with SRE, Compute, and Engineering teams, you will lead root‑cause analysis, identify permanent solutions to recurring issues, and translate operational learnings into stronger monitoring, processes, and runbooks.

You will also play an important role during major incidents and post‑incident reviews, while mentoring L1/L2 engineers and improving the quality of technical documentation and escalation practices. The position participates in an on‑call escalation rotation.

Requirements

You should have at least 6 years of experience in cloud infrastructure technical support, escalation engineering, production operations, or an SRE‑adjacent role, with strong hands‑on expertise across Linux, networking, and cloud infrastructure. Experience supporting GPU infrastructure, CUDA, HPC, or large‑scale compute environments will be particularly relevant.

You will bring a strong track record of diagnosing difficult production issues, performing structured root‑cause analysis, and collaborating effectively with engineering teams to drive issues through to resolution.

Strong written and verbal communication skills, with the ability to communicate clearly across technical teams, are important, along with comfort handling production incidents and participating in an on‑call environment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU & Cloud Infrastructure Support Engineer
Senior GPU & Cloud Infrastructure Support Engineer

Kerry Consulting • Singapore

On-site
SGD 120,000 - 180,000
Data Centre Operations Engineer (Singapore, Singapore)
Data Centre Operations Engineer (Singapore, Singapore)

Singtel • Singapore

On-site
Confidential
Health and wellness benefits
Training and development programs
Internal mobility opportunities
DevOps Engineer, GPUaaS
DevOps Engineer, GPUaaS

Singapore Telecommunications Limited • Singapore

On-site
SGD 120,000 - 180,000
Senior AI/HPC Compute Infrastructure Engineer
Senior AI/HPC Compute Infrastructure Engineer

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 190,000
Senior AI Infrastructure Support Engineer
Senior AI Infrastructure Support Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Infrastructure Reliability Engineer
Senior AI Infrastructure Reliability Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Data Centre Operations Engineer
Data Centre Operations Engineer

Singapore Telecommunications Limited • Singapore

On-site
SGD 60,000 - 90,000
Health benefits
Training & development
Internal mobility
Network Engineer, GPUaaS (Singapore, Singapore)
Network Engineer, GPUaaS (Singapore, Singapore)

Singtel • Singapore

On-site
Confidential
DevOps Engineer, GPUaaS
DevOps Engineer, GPUaaS

Singtel Group • Singapore

On-site
SGD 120,000 - 170,000
6723 - GPU Infrastructure Engineer | Up to $7K | Kaki Bukit | NVIDIA, CUDA & HPC
6723 - GPU Infrastructure Engineer | Up to $7K | Kaki Bukit | NVIDIA, CUDA & HPC

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000