NOC Engineer, AI Infrastructure (L1 / L2 / L3)

Lancesoft

Dadri

On-site

INR 600,000 - 900,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Job summary

LanceSoft is building a NOC to monitor and support AI and GPU-accelerated infrastructure for enterprise clients, including DGX/HGX systems and GPU clusters. You will be the front line for keeping these environments healthy, performant and available.

Key responsibilities include monitoring GPU health, triaging alerts, following runbooks, and collaborating with data center teams. The role offers opportunities across L1–L3 levels with growing scope in automation and incident leadership.

Qualifications

  • Bachelor's degree in Computer Science, IT, Electronics or a related field.
  • Strong written and verbal communication in English.
  • Willingness to work rotational shifts, weekends and on-call schedules.

Responsibilities

  • Monitor GPU clusters, servers, networks, storage and AI platforms through the NOC dashboards.
  • Respond to alerts within SLA and perform first-level triage.
  • Log, categorize, prioritize and track incidents in ITSM tools.
  • Follow runbooks and escalation matrices and keep handover notes accurate.
  • Communicate incident status clearly to internal teams and client stakeholders.
  • Contribute to knowledge base articles and runbook improvements.
  • Lead or participate in incident bridges and post-incident reviews.

Skills

Linux administration
Python
Bash
NVIDIA DGX
NVIDIA HGX
NVIDIA DCGM
NVIDIA GPU Operator
NVIDIA NGC containers
InfiniBand
NCCL
Kubernetes
Slurm
Docker
containerd

Education

Bachelor's degree in Computer Science, IT, Electronics or related field

Tools

Docker
containerd
NVIDIA Base Command Manager
DCGM
Prometheus
Grafana
ServiceNow
Jira

Job description

About the Role

LanceSoft is building a NOC that monitors and supports AI and GPU-accelerated infrastructure for enterprise clients. This includes NVIDIA DGX and HGX systems, GPU clusters, high-speed networking, and AI workload platforms. You will be the front line for keeping these environments healthy, performant and available.

Required Certifications by Experience Level
L1 Entry Level
  • Required: NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • Preferred: Candidates currently working toward the NVIDIA-Certified Professional: AI Operations (NCP-AIO) certification.
L2 Intermediate Level
  • Required: NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • Preferred: NVIDIA-Certified Professional: AI Operations (NCP-AIO)
L3 Senior Level
  • Required: NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO) and NVIDIA-Certified Professional: AI Operations (NCP-AIO)
  • Preferred: Additional certifications in cloud computing, networking, or related infrastructure technologies.
Common Responsibilities (All Levels)
  • Monitor GPU clusters, servers, networks, storage and AI platforms through the NOC dashboards, and respond to alerts within SLA.
  • Log, categorize, prioritize and track incidents in the ITSM tool (ServiceNow, Jira or similar).
  • Follow runbooks and escalation matrices, and keep shift handover notes accurate.
  • Communicate incident status clearly to internal teams and client stakeholders.
  • Contribute to knowledge base articles and runbook improvements.
  • Follow ITIL-aligned incident, problem and change practices.
L1 NOC Engineer: Monitoring and First Response

Experience: 0-2 years in NOC, IT support or data center operations

Responsibilities
  • 24x7 monitoring of GPU health, utilization, temperature, power and node availability.
  • First-level triage and validation of alerts. Perform basic checks such as node reachability, service status and job queue status.
  • Execute documented remediation steps, such as restarting services or draining and returning nodes, following runbooks.
  • Escalate unresolved or complex incidents to L2 with complete diagnostic details.
  • Coordinate with data center or remote-hands teams on hardware tickets.
Skills
  • Foundational understanding of AI infrastructure: GPUs, DGX and HGX systems, NVLink, InfiniBand and Ethernet fabrics.
  • Basic Linux command line, networking fundamentals (TCP/IP, DNS, VLANs) and ticketing tools.
  • Familiarity with nvidia-smi and basic monitoring dashboards.
L2 NOC Engineer: Investigation and Resolution

Experience: 3-5 years in infrastructure operations, with exposure to GPU or HPC environments

Responsibilities
  • Investigate and resolve escalated incidents, including GPU faults, XID errors, driver and firmware issues, node failures and job scheduling problems.
  • Administer and troubleshoot workload schedulers and orchestration (Slurm, Kubernetes with the NVIDIA GPU Operator).
  • Analyze performance issues across compute, network (InfiniBand/RoCE) and storage, and perform root cause analysis.
  • Manage cluster operations through NVIDIA Base Command Manager or equivalent tooling.
  • Configure and tune monitoring and alerting (DCGM, Prometheus, Grafana) to reduce noise and improve detection.
  • Execute changes, patching, firmware and driver updates under change control.
  • Mentor L1 engineers and review the quality of escalations.
Skills
  • Strong Linux administration and shell scripting (Bash, Python).
  • Hands-on experience with GPU diagnostics, DCGM, container runtimes (Docker, containerd) and NVIDIA NGC containers.
  • Working knowledge of InfiniBand and NCCL, and of storage for AI (parallel or NFS-based).
L3 NOC Engineer: Senior and Escalation Lead

Experience: 6-10 years in infrastructure or platform operations, including 2+ years on AI, GPU or HPC platforms

Responsibilities
  • Act as the final technical escalation point for critical incidents and major outages. Lead incident bridges and post-incident reviews.
  • Perform deep diagnostics on multi-node training and inference performance, fabric congestion, NVLink and InfiniBand faults, and GPU memory or ECC errors.
  • Design and improve NOC monitoring architecture, observability, alert thresholds and SLO/SLA reporting for AI infrastructure.
  • Automate operations using Python, Ansible and Terraform. Build self-healing workflows and reduce manual toil.
  • Own problem management, capacity planning and lifecycle management (firmware, driver and CUDA compatibility matrices).
  • Engage with NVIDIA and OEM support (Dell, HPE, Supermicro and others), and manage vendor cases through resolution.
  • Define runbooks, standards and training plans for L1 and L2, and contribute to client solution reviews and pre-sales technical input where needed.
Skills
  • Expert-level understanding of the NVIDIA AI stack: DGX, Base Command, DCGM, NCCL, GPU Operator, NGC, and NVIDIA AI Enterprise.
  • Advanced Kubernetes and Slurm operations, and InfiniBand fabric management (UFM).
  • Strong scripting and automation, with an observability and incident-command mindset.
Common Qualifications
  • Bachelor's degree in Computer Science, IT, Electronics or a related field (or equivalent experience).
  • The certifications listed for the level above (validity must be current).
  • Willingness to work rotational shifts, weekends and on-call schedules.
  • Strong written and verbal communication in English.
Nice to Have
  • ITIL 4 Foundation
  • Cloud certifications (AWS, Azure or GCP)
  • CCNA or CCNP
  • RHCSA or CKA
  • Experience with air-gapped or on-prem AI environments
  • Experience with the ServiceNow, Zabbix, SolarWinds or Nagios toolsets
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Platform Engineer
GPU Platform Engineer

Lancesoft • Dadri

On-site
INR 3,500,000 - 6,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
AI Engineer
AI Engineer

STCO India • Hyderabad

On-site
INR 800,000 - 1,200,000
NCX Senior Engineer
NCX Senior Engineer

NVIDIA • Maharashtra

On-site
INR 3,000,000 - 5,500,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000
NCX Senior Engineer
NCX Senior Engineer

NVIDIA • Hyderabad

On-site
INR 4,000,000 - 6,500,000
NCX Senior Engineer
NCX Senior Engineer

NVIDIA Gruppe • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Software Solutions Engineer
Software Solutions Engineer

NVIDIA Corporation • Pune District

On-site
INR 1,800,000 - 2,600,000
On-call weekend rotation
Software Solutions Engineer
Software Solutions Engineer

NVIDIA • Pune District

On-site
INR 1,800,000 - 2,800,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Corporation • Mumbai

On-site
INR 3,500,000 - 5,000,000