L3 Data Center Engineer

Lintasarta

Jakarta Pusat

On-site

IDR 140,000,000 - 320,000,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lintasarta seeks an experienced L3 Engineer / GPU Infrastructure Specialist to act as the highest technical authority within the managed service delivery team. You will engage when firmware-level analysis or vendor tooling is required, with on-call availability and onsite mobilisation as needed.

Responsibilities include leading OEM escalations, performing advanced NVIDIA InfiniBand and liquid cooling troubleshooting, and producing formal RCA reports.

Qualifications

  • Min. 7 years in Data Center / infrastructure with hands-on GPU server/HPC expertise at firmware and OEM level.
  • Advanced GPU diagnostics including Nvidia-smi, lspci, PCIe/NVLink topology, and ECC/XID error tracking.
  • Experience with NVIDIA InfiniBand networking, fabric troubleshooting, and escalation processes with vendors.

Responsibilities

  • Handle incidents beyond L2 resolution, including firmware-level failure analysis and OEM tooling.
  • Own OEM escalation lifecycle and interface with NVIDIA and server/network OEM teams through to resolution.
  • Escalate to Principal Expert for manufacturer-level intervention or multi-vendor failures.
  • Act as technical reference for L1/L2 teams in real-time diagnostics and cross-stack resolution strategies.
  • Lead advanced troubleshooting of NVIDIA InfiniBand fabric and liquid cooling system root-cause analysis.
  • Attend governance reviews to advise on failure trends and infrastructure risk.
  • Contribute to knowledge base with escalation playbooks and training for junior staff.

Skills

GPU servers/HPC
Firmware & OEM level diagnostics
NVIDIA InfiniBand networking
Fabric troubleshooting
NVIDIA diagnostics
OEM escalation leadership

Tools

OEM diagnostic tooling

Job description

The L3 Engineer / GPU Infrastructure Specialist is the highest technical authority within the managed service delivery team, engaged when L2 has exhausted all standard diagnostic and replacement procedures and the incident requires firmware-level analysis, OEM-specific tooling, or direct vendor engineering support. Operating on a 24×7 on-call basis with the ability to mobilise onsite when required, L3 specialists serve simultaneously as deep technical resolvers, OEM relationship owners, and knowledge anchors for the broader delivery team.

Responsibilities:
  • Handles incidents beyond L2 resolution scope, including firmware-level failure analysis, NVLink/PCIe fabric anomalies, deep XID and ECC error investigations, and the application of OEM-specific diagnostic tooling unavailable at lower tiers; produces formal RCA and failure reports for every L3 engagement.
  • Owns the complete OEM escalation lifecycle, from defining trigger conditions and assembling evidence packages, to direct interface with NVIDIA, server OEM, and network OEM engineering teams, through to resolution tracking and formal closure.
  • Escalates to the Principal Expert when incidents exceed L3 resolution capability, including cases requiring manufacturer-level engineering intervention, multi-vendor cross-stack failures, or issues with no established resolution precedent.
  • Serves as the authoritative technical reference point for L1 and L2 personnel throughout incident handling, providing real-time diagnostic guidance and directing cross-stack resolution strategy.
  • Leads advanced troubleshooting of the NVIDIA InfiniBand backend fabric (fabric health, NCCL collective performance, network congestion) and root-cause analysis of liquid-cooling system failures at CDU and secondary loop level.
  • Attends regular governance review meetings to advise on failure trends, troubleshooting strategy, and infrastructure risk.
  • Contributes to the operational knowledge base through escalation playbooks, known-issue documentation, and structured competency uplift of the L1/L2 team via the shadow/pairing model.
Qualifications & Skills:
  • Min. 7 years in Data Center / infrastructure, with deep hands-on expertise on GPU servers / HPC clusters at firmware and OEM level, beyond day-to-day operations.
  • Advanced GPU diagnostics (Nvidia-smi/lspci, PCIe/NVLink link & topology, thermal/power, ECC & XID error tracking).
  • NVIDIA InfiniBand networking, fabric troubleshooting, link/subnet manager, OEM escalation.
  • Liquid cooling systems (CDU & secondary loop) to root-cause analysis level.
  • Demonstrated end-to-end OEM escalation leadership.
  • Manufacturer-level capability for complex, cross-stack infrastructure issues.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Engineer (L2)
Data Center Engineer (L2)

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Senior GPU Infra & OEM Escalation Engineer
Senior GPU Infra & OEM Escalation Engineer

Lintasarta • Jakarta Pusat

On-site
IDR 140,000,000 - 320,000,000
GPU Data Center Engineer — 24/7 Infra Fault Specialist
GPU Data Center Engineer — 24/7 Infra Fault Specialist

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Data Centre Field Operations Engineer
Data Centre Field Operations Engineer

Nava • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000
Data Center Field Ops Lead - GPU/HPC & Vendor Governance
Data Center Field Ops Lead - GPU/HPC & Vendor Governance

Nava • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000
Cloud & Data Center Engineer (GPU & OpenStack)
Cloud & Data Center Engineer (GPU & OpenStack)

Lintasarta • Kota Medan ᯔᯩᯑᯉ᯲

On-site
IDR 120,000,000 - 200,000,000
Data Center Infrastructure Engineer - GPU & OpenStack
Data Center Infrastructure Engineer - GPU & OpenStack

Lintasarta • Kota Semarang

On-site
IDR 133,920,000 - 267,840,000
Data Center Civil Engineer - On-Site Project Coordinator
Data Center Civil Engineer - On-Site Project Coordinator

Pt Trimitra Putra Mandiri • Jawa Barat

On-site
IDR 72,000,000 - 108,000,000
Storage and Data Service Engineer (HPC)
Storage and Data Service Engineer (HPC)

Nityo Infotech Indonesia • Batang

On-site
IDR 150,000,000 - 210,000,000
Senior Server Administrator
Senior Server Administrator

Digital Edge Data Center • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000