AI Infrastructure Support Engineer

NSCALE OPERATIONS APAC PTE. LTD.

Singapore

On-site

SGD 65,000 - 95,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NSCALE OPERATIONS APAC PTE. LTD. is seeking an experienced infrastructure support engineer to join our on-site Singapore team. You will handle tickets, triage GPU nodes, perform hardware diagnostics, and collaborate with Engineering during incidents or changes.

Responsibilities include runbook-led diagnostics, evidence collection for RMAs, observing and improving dashboards, and contributing to automation. Travel to customer sites may be required as part of deployments and support.

Qualifications

  • 3+ years in infrastructure support or support engineering in SLA-driven environments.
  • Clear written notes and reliable follow-through.
  • Hands-on GPU infrastructure experience with nvidia-smi or similar.
  • Strong Linux CLI skills and basic networking tools.
  • Knowledge of IP addressing, subnets and VLANs.
  • Experience with ITSM processes and ticketing.

Responsibilities

  • Handle day-to-day tickets and alerts; escalate early with guidance.
  • GPU node triage and hardware troubleshooting; prepare vendor RMA evidence.
  • Run fabric and link diagnostics; capture evidence for handover.
  • Assist storage and data-path investigations on high-performance platforms.
  • Record and resolve tickets with clear customer communications.
  • Participate in on-call and travel to customer locations as required.

Skills

Infrastructure support
Communication
GPU hardware troubleshooting
BMC management
Linux CLI
Networking basics
ITSM processes
Observability basics
Scripting: Bash
Scripting: Python
Git version control
Platform & DC fundamentals
Adaptability
Growth mindset

Tools

nvidia-smi
NetBox

Job description

What You'll be Doing
  • Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it
  • Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA
  • Run fabric and link diagnostics following established runbooks (mlxlink orequivalent), capture evidence accurately, and escape with a handover that lets the next engineer continue without starting from scratch
  • Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis
  • Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review
  • Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels
  • Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover
  • Participate in changes under peer review, learning risk assessment and backout practices in live customer environments
  • Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns)
  • Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes
  • Be the escalation point for onsite DC Operations staff; coordinate smart- hands tasks within your scope
  • Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed
  • Share knowledge by documenting steps you've validated and contributing to training materials.
  • Take part in incident reviews as a contributor and help track preventative follow-ups in your scope
  • Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed
  • Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required.
About You
  • Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services)
  • Communication. Clear written notes, concise updates, and reliable follow- through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately
  • GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via
  • BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal
  • Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalation
  • Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role
  • Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation
  • Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review
  • Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control
  • Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background
  • Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior
  • Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Support Engineer
Senior AI Infrastructure Support Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Network Operations Engineer
Network Operations Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Infrastructure Lead (Network & Systems)
Infrastructure Lead (Network & Systems)

NEUTRON PTE. LTD. • Singapore

On-site
SGD 180,000 - 240,000
Data Center Operations Engineer
Data Center Operations Engineer

Singtel • Singapore

On-site
Confidential
Infrastructure Support Engineer (SG)
Infrastructure Support Engineer (SG)

NTC Integration Pte Ltd • Singapore

On-site
SGD 42,000 - 65,000
Onsite Support Engineer
Onsite Support Engineer

TEKISHUB CONSULTING SERVICES PTE. LTD. • Singapore

On-site
SGD 52,000 - 76,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
FIS Line Support Engineer
FIS Line Support Engineer

ADVANCED E-SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 48,000 - 72,000
Information Technology - Lead Systems Engineer, IT Network Comms (Core DC)
Information Technology - Lead Systems Engineer, IT Network Comms (Core DC)

SINGAPORE AIRLINES LIMITED • Singapore

On-site
SGD 120,000 - 180,000
Infrastructure Support Engineer, Linux (Contract)
Infrastructure Support Engineer, Linux (Contract)

GMP TECHNOLOGIES (S) PTE LTD • Singapore

On-site
SGD 70,000 - 110,000