AI Infrastructure Support Engineer - APAC (GPUs)

Nscale

Singapore

On-site

SGD 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale is recruiting an AI Infrastructure Support Engineer for APAC to handle day‑to‑day tickets, hardware triage, and incident response. You’ll work with Engineering and perform GPU node diagnostics, run network and storage checks, and document steps for efficient handovers.

The role requires 3+ years in infrastructure support, strong Linux CLI skills, and familiarity with ITSM processes. Occasional on‑call and travel to customer sites are expected.

Qualifications

  • 3+ years in infrastructure support or support engineering in SLA‑driven customer environments.
  • Clear written notes, concise updates, and reliable handovers to customers and teams.
  • GPU and hardware troubleshooting using nvidia-smi or similar diagnostics.
  • Solid CLI skills with Linux systemd, filesystems, permissions, and networking.
  • Experience with ITSM processes (ITIL or similar) and ticketing discipline.

Responsibilities

  • Handle day‑to‑day tickets and alerts; escalate incidents with guidance.
  • Triage GPU nodes; interpret logs and perform hardware remediation.
  • Run fabric and link diagnostics; capture evidence for handover.
  • Investigate storage and data‑path issues; gather evidence for diagnosis.
  • Record and manage tickets; communicate with customers clearly.
  • Participate in on‑call requirements and travel for deployments.

Job description

About Nscale
Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance infrastructure for AI start‑ups and large enterprise customers. Nscale enables AI‑focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability and rapid response to customer tickets. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

AI Infrastructure Support Engineer – APAC
What You’ll Be Doing
  • Join the Support duty rotation and handle day‑to‑day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it
  • Perform GPU node triage and hardware troubleshooting: interpret nvidia‑smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA
  • Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and elevate with a handover that lets the next engineer continue without starting from scratch
  • Assist with storage and data‑path investigations (mounts, connectivity, client‑side symptoms) on high‑performance platforms, gathering evidence for Senior or Engineering‑led diagnosis
  • Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review
  • Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels
  • Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover
  • Participate in changes under peer review, learning risk assessment and backout practices in live customer environments
  • Help maintain source‑of‑truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns)
  • Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes
  • Be the escalation point for onsite DC Operations staff; coordinate smart‑hands tasks within your scope
  • Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed
  • Share knowledge by documenting steps you’ve validated and contributing to training materials. Shadow Seniors during complex work to build capability
  • Take part in incident reviews as a contributor and help track preventative follow‑ups in your scope
  • Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed
  • Participate in on‑call and out‑of‑hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required
About You
  • 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA‑driven, customer‑facing environments (cloud, data centre, or managed services)
  • Clear written notes, concise updates, and reliable follow‑through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately
  • GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands‑on with nvidia‑smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers — reseating components, swap testing, working via BMC/out‑of‑band management — through to preparing RMA evidence. A strong technical base here is required, not a learning goal
  • Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to elevate
  • Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high‑performance east‑west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role
  • Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation
  • Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review
  • Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control
  • Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background
  • Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior
  • Able to work in a fast‑moving environment with evolving processes, participate in on‑call after onboarding, and travel when needed
Nice to Have
  • High‑performance fabrics and GPU‑HPC. Exposure to RDMA/InfiniBand, link‑level diagnostics (mlxlink, ibdiagnet, or equivalent), NCCL‑based troubleshooting, or NVLink concepts
  • High‑performance storage. Exposure to VAST or comparable AI‑optimised storage platforms, Ceph, or NFS at scale, including basic storage‑network troubleshooting
  • OpenStack and fleet operations tooling. Familiarity with OpenStack troubleshooting flows, or fleet‑scale tooling for provisioning and health (MAAS, NetBox, Redfish, or similar)
  • Understanding of core concepts (nodes, pods, services, logs) and basic troubleshooting via runbooks. Helpful context for our platform, though not the core of this role
  • Automation and access tooling. Experience with Ansible or Terraform, CI/CD participation (GitHub Actions or similar), or access and security tooling such as Teleport or Vault
  • Progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time
In All We Do, Our Core Values Guide Us

Relentless Innovation – At Nscale, we constantly push the boundaries of innovation…

Ownership and Accountability – Every Nscaler is fully accountable…

Openness and Transparency – We believe trust and transparency are key…

Customer‑Centric Focus – Our customers are central…

Sustainability – We are dedicated to considering the long‑term environmental…

Full‑Speed Collaboration – Collaboration at Nscale is fast, efficient, and respectful…

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio‑economic backgrounds. If there’s anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Deployment Engineering (GPU Systems & Networking)
Senior Manager, Deployment Engineering (GPU Systems & Networking)

Nscale • Singapore

On-site
SGD 180,000 - 240,000
Solutions Architect (AI GPU Infrastructure / Data Centre Architecture)
Solutions Architect (AI GPU Infrastructure / Data Centre Architecture)

Nscale • Singapore

On-site
SGD 180,000 - 240,000
Principal Systems Engineer - APAC
Principal Systems Engineer - APAC

Nscale • Singapore

On-site
SGD 180,000 - 240,000
SVP, Data Centre (Design & Construction) - APAC
SVP, Data Centre (Design & Construction) - APAC

Nscale • Singapore

On-site
SGD 300,000 - 600,000
Principal Network Engineer – APAC
Principal Network Engineer – APAC

Nscale • Singapore

On-site
SGD 160,000 - 240,000
Data Engineer (Forward Deployed)
Data Engineer (Forward Deployed)

Nscale • Singapore

On-site
SGD 70,000 - 100,000
Highly competitive package (base + equity)
Dynamic progression plan tailored to your ambitions
Human-First Flexibility
SVP, Data Centre (Design & Construction) - APAC Singapore
SVP, Data Centre (Design & Construction) - APAC Singapore

Nscale • Singapore

On-site
SGD 450,000 - 750,000
GPU AI Infrastructure Support Engineer
GPU AI Infrastructure Support Engineer

Nscale • Singapore

On-site
SGD 70,000 - 110,000
Senior Manager – Site Selection & Capacity Program Management (For Japan/Australia markets)
Senior Manager – Site Selection & Capacity Program Management (For Japan/Australia markets)

Nscale • Singapore

On-site
SGD 232,000 - 362,000
Solutions Architect, Data Center Infrastructure
Solutions Architect, Data Center Infrastructure

NVIDIA Corporation • Singapore

On-site
SGD 90,000 - 150,000