Data Center Engineer (L1)

Lintasarta

Jakarta Pusat

On-site

IDR 89,280,000 - 156,240,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Lintasarta seeks an L1 Engineer to operate across two on-site sub-functions: Surveillance (NOC monitoring) and Data Center (on-site response) within a 24×7 coverage model.

You will monitor GPU infrastructure dashboards, classify alerts, create ITSM tickets, and escalate as needed, ensuring rapid incident response and SLA adherence.

Qualifications

  • Min. 1–2 years in Data Center / NOC / IT Operations.
  • Basic TCP/IP networking.
  • Proficient with monitoring tools (Grafana / Zabbix / DCIM / NMS) & ITSM ticketing systems.
  • Solid understanding of SLA concepts, incident prioritization, and escalation workflows.

Responsibilities

  • Maintains continuous 24×7 monitoring of GPU infrastructure dashboards, including compute health, GPU utilization, fabric and network status, power/cooling, and environment.
  • Classifies alarms by severity (P1–P4) and dispatches through the escalation workflow with SLA tracking.
  • Creates accurate incident tickets in ITSM and routes to the correct tier (L1, L2, L3).
  • Provides P1 status updates every 30 minutes and monitors SLA countdown, escalating at risk of breach.

Skills

NOC Operations
Data Center
Incident Escalation
SLA Awareness

Tools

Grafana
Zabbix
DCIM
NMS

Job description

The L1 Engineer forms the foundational layer of the tiered support model, operating across two complementary sub-functions within the same 24×7 onsite coverage model: Surveillance (NOC-based continuous monitoring) and Data Center (onsite physical response). Together, these two sub-functions ensure that every infrastructure event, whether detected remotely through dashboards or observed directly on the data center floor, is captured, classified, and escalated with precision and speed. L1 is the first point of contact for all incidents and the initiator of the escalation path to L2 and L3.


Responsibilities:


  • Maintains continuous 24×7 monitoring of all GPU infrastructure dashboards, covering compute health, GPU utilization, fabric and network status, power and cooling parameters, and environmental conditions, using platforms including DCGM, NetQ, UFM, Grafana, and ServiceNow.

  • Classifies all alarms by severity (P1–P4), validates against false-positive filters, and dispatches through the appropriate escalation workflow with full SLA tracking.

  • Creates accurate, complete incident tickets in the ITSM platform and dispatches to the appropriate tier (L1 DC for physical check, L2 for technical diagnosis, or L3 for complex escalation).

  • Provides P1 status updates every 30 minutes until resolution; monitors SLA countdown for all active tickets and proactively escalates tickets at risk of breach.

  • Conducts daily synthetic health checks: canary jobs, NCCL bandwidth tests, fabric monitoring, and log pipeline health verification.

  • Conducts regular physical walkthroughs and rack inspections, verifying LED indicators, cabling integrity, power supply status, and the physical condition of GPU servers, NVLink switches, and CDUs.

  • Responds to ticket dispatch from L1 Surveillance for direct on-floor physical checks, and performs basic hardware verification and initial corrective actions (e.g., reboot/power-cycle via BMC) before escalating to L2.

  • Executes Emergency Response Procedures (EPO or Loop Isolation protocols) within 15 minutes of alert confirmation for P1 environmental incidents including cooling failures, power irregularities, or liquid coolant leaks.

  • Monitors liquid-cooling system parameters (CDU and secondary loop) and coordinates with the DC facilities team on environmental anomalies.

  • Supports preventive maintenance activities: rack deep cleaning (quarterly), node health sweep (monthly), and cold-spare rotation.

  • Executes a structured shift handover at every transition, ensuring full situational awareness of active incidents, open tickets, pending dispatches, and infrastructure anomalies is formally transferred to the incoming shift with zero information loss.


Qualifications & Skills :


  • Min. 1–2 years in Data Center / NOC / IT Operations.

  • Basic TCP/IP networking.

  • Proficient with monitoring tools (Grafana / Zabbix / DCIM / NMS) & ITSM ticketing systems.

  • Solid understanding of SLA concepts, incident prioritization, and escalation workflows.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Center Engineer (L2)
Data Center Engineer (L2)

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
L1 Data Centre Operator
L1 Data Centre Operator

PT Leap Digital Indonesia • Jakarta Pusat

On-site
IDR 133,920,000 - 234,360,000
L3 Data Center Engineer
L3 Data Center Engineer

Lintasarta • Jakarta Pusat

On-site
IDR 140,000,000 - 320,000,000
Data Centre Field Operations Engineer
Data Centre Field Operations Engineer

Nava • Jakarta Pusat

On-site
IDR 350,000,000 - 550,000,000
GPU Data Center Engineer — 24/7 Infra Fault Specialist
GPU Data Center Engineer — 24/7 Infra Fault Specialist

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Data Center Operations Engineer
Data Center Operations Engineer

BYD • Subang

On-site
IDR 100,440,000 - 156,240,000
Data Center Engineer
Data Center Engineer

PT. Permodalan Nasional Madani (Persero) • Jakarta Selatan

On-site
IDR 167,400,000 - 256,680,000
Service Delivery Field Support Engineer (L1)
Service Delivery Field Support Engineer (L1)

NTT Limited • Daerah Khusus Ibukota Jakarta

Hybrid
Senior GPU Infra & OEM Escalation Engineer
Senior GPU Infra & OEM Escalation Engineer

Lintasarta • Jakarta Pusat

On-site
IDR 140,000,000 - 320,000,000
Data Center Engineer/Lead
Data Center Engineer/Lead

Atomic Recruitment SEA • Jakarta Pusat

On-site
IDR 111,600,000 - 189,720,000