Data Center Engineer (L2)

Lintasarta

Jakarta Pusat

On-site

IDR 279,000,000 - 502,200,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Lintasarta is seeking an L2 Engineer to own and resolve complex GPU and infrastructure incidents in a large-scale GPU cloud environment. You will diagnose hardware faults across GPU servers, cooling systems, and high-speed networks, while maintaining thorough RCA documentation and keeping firmware baselines current.

The role requires hands-on experience with GPU diagnostics, NVLink, InfiniBand networking, and enterprise storage fault isolation.

Qualifications

  • Min. 3–5 years in Data Center / infrastructure operations with hands-on GPU server experience.
  • In-depth GPU diagnostics, PCIe/NVLink, thermal and power monitoring.
  • Rack-scale GB200 NVL72 and liquid-cooling operation experience.
  • NVIDIA InfiniBand networking and enterprise storage fault isolation.
  • Strong RCA documentation and technical writing skills.

Responsibilities

  • Performs in-depth fault diagnosis and hardware-level resolution across the full infrastructure stack, including GPU server faults, liquid-cooling anomalies, and network issues.
  • Authors Root Cause Analysis (RCA) documentation for every incident.
  • Governs all infrastructure changes through the CAB, coordinating changes and maintenance windows.
  • Initiates formal escalation to L3/OEM with complete evidence packages when standard procedures are exhausted.
  • Maintains firmware and driver baseline compliance and monitors CVE advisories and OEM bulletins.
  • Supports knowledge transfer through a shadow/pairing model with IOH's embedded operational counterparts.

Skills

GPU servers
HPC clusters
NVIDIA InfiniBand
Liquid-cooling systems
NVLink
GB200 NVL72
Hardware troubleshooting
Root Cause Analysis
Technical writing
Change management

Tools

nvidia-smi
lspci
DCGM

Job description

The L2 Engineer is the technical resolution authority for all infrastructure faults that exceed L1's first-response capability. Operating onsite 24×7, L2 engineers own each incident ticket from the point of escalation through to hardware-level resolution or formal handoff to L3, maintaining ticket ownership throughout the entire escalation lifecycle. The role demands deep, hands-on expertise across GPU compute, liquid-cooling systems, high-speed network fabric, and server infrastructure in a large-scale, mission-critical GPU cloud environment.

Key Responsibilities:
  • Performs in-depth fault diagnosis and hardware-level resolution across the full infrastructure stack, including GPU server faults (card detection, PCIe/NVLink link status, thermal and power anomalies, ECC error tracking), liquid-cooling system anomalies (CDU and secondary loop: inlet/outlet temperature, flow rate, differential pressure), NVIDIA InfiniBand and converged network faults, and enterprise server and storage issues.
  • Authors Root Cause Analysis (RCA) documentation for every incident, capturing fault timeline, diagnostic findings, actions taken, and preventive recommendations.
  • Governs all infrastructure changes through the CAB, accountable for change request submission, approval coordination, maintenance window execution, rollback planning, and post-change verification.
  • Initiates formal escalation to L3/OEM with a complete evidence package (diagnostic logs, DCGM output, hardware status reports, incident timelines) when standard diagnostic and replacement procedures are exhausted, while retaining ticket ownership until full resolution.
  • Maintains firmware and driver baseline compliance across all in-scope infrastructure, monitoring CVE advisories and OEM bulletins, executing baseline drift corrections within approved maintenance windows, with minimum 90-day configuration backup retention.
  • Supports the knowledge transfer programme through a shadow/pairing model with IOH's embedded operational counterparts from service commencement.
Qualifications & Skills:
  • Min. 3–5 years in Data Center / infrastructure operations, with direct hands-on experience on GPU servers / HPC clusters (not general servers only).
  • In-depth GPU diagnostics (nvidia-smi/lspci, PCIe/NVLink, thermal & power, ECC error tracking).
  • Rack-scale GB200 NVL72 (Schedule A - Voltage).
  • Liquid-cooling systems (CDU & secondary loop), operation, monitoring, and anomaly handling (Schedule A - Voltage).
  • NVIDIA InfiniBand + converged/management networking.
  • Enterprise server & storage: fault isolation & vendor escalation.
  • Strong RCA documentation and technical writing skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Data Center Engineer — 24/7 Infra Fault Specialist
GPU Data Center Engineer — 24/7 Infra Fault Specialist

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Data Centre Field Operations Engineer
Data Centre Field Operations Engineer

Nava • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000
Cloud & Data Center Engineer (GPU & OpenStack)
Cloud & Data Center Engineer (GPU & OpenStack)

Lintasarta • Kota Medan ᯔᯩᯑᯉ᯲

On-site
IDR 120,000,000 - 200,000,000
Data Center Field Ops Lead - GPU/HPC & Vendor Governance
Data Center Field Ops Lead - GPU/HPC & Vendor Governance

Nava • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000
Data Center Infrastructure Engineer - GPU & OpenStack
Data Center Infrastructure Engineer - GPU & OpenStack

Lintasarta • Kota Semarang

On-site
IDR 133,920,000 - 267,840,000
Data Center Engineer
Data Center Engineer

PT. Permodalan Nasional Madani (Persero) • Jakarta Selatan

On-site
IDR 167,400,000 - 256,680,000
Data Center Engineer
Data Center Engineer

Esha Parama Technology • Daerah Khusus Ibukota Jakarta

On-site
IDR 180,000,000 - 280,000,000
Cloud Engineer L2
Cloud Engineer L2

PT. Mastersystem Infotama • Jawa Barat

On-site
IDR 267,840,000 - 468,720,000
Storage and Data Service Engineer (HPC)
Storage and Data Service Engineer (HPC)

Nityo Infotech Indonesia • Batang

On-site
IDR 150,000,000 - 210,000,000
Cloud Engineer
Cloud Engineer

PT Datacomm Diangraha • Jakarta Pusat

On-site
IDR 180,000,000 - 300,000,000