About the Role
LanceSoft is building a NOC that monitors and supports AI and GPU-accelerated infrastructure for enterprise clients. This includes NVIDIA DGX and HGX systems, GPU clusters, high-speed networking, and AI workload platforms. You will be the front line for keeping these environments healthy, performant and available.
Required Certifications by Experience Level
L1 Entry Level
- Required: NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
- Preferred: Candidates currently working toward the NVIDIA-Certified Professional: AI Operations (NCP-AIO) certification.
L2 Intermediate Level
- Required: NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
- Preferred: NVIDIA-Certified Professional: AI Operations (NCP-AIO)
L3 Senior Level
- Required: NVIDIA-Certified Associate: AI Infrastructure & Operations (NCA-AIIO) and NVIDIA-Certified Professional: AI Operations (NCP-AIO)
- Preferred: Additional certifications in cloud computing, networking, or related infrastructure technologies.
Common Responsibilities (All Levels)
- Monitor GPU clusters, servers, networks, storage and AI platforms through the NOC dashboards, and respond to alerts within SLA.
- Log, categorize, prioritize and track incidents in the ITSM tool (ServiceNow, Jira or similar).
- Follow runbooks and escalation matrices, and keep shift handover notes accurate.
- Communicate incident status clearly to internal teams and client stakeholders.
- Contribute to knowledge base articles and runbook improvements.
- Follow ITIL-aligned incident, problem and change practices.
L1 NOC Engineer: Monitoring and First Response
Experience: 0-2 years in NOC, IT support or data center operations
Responsibilities
- 24x7 monitoring of GPU health, utilization, temperature, power and node availability.
- First-level triage and validation of alerts. Perform basic checks such as node reachability, service status and job queue status.
- Execute documented remediation steps, such as restarting services or draining and returning nodes, following runbooks.
- Escalate unresolved or complex incidents to L2 with complete diagnostic details.
- Coordinate with data center or remote-hands teams on hardware tickets.
Skills
- Foundational understanding of AI infrastructure: GPUs, DGX and HGX systems, NVLink, InfiniBand and Ethernet fabrics.
- Basic Linux command line, networking fundamentals (TCP/IP, DNS, VLANs) and ticketing tools.
- Familiarity with nvidia-smi and basic monitoring dashboards.
L2 NOC Engineer: Investigation and Resolution
Experience: 3-5 years in infrastructure operations, with exposure to GPU or HPC environments
Responsibilities
- Investigate and resolve escalated incidents, including GPU faults, XID errors, driver and firmware issues, node failures and job scheduling problems.
- Administer and troubleshoot workload schedulers and orchestration (Slurm, Kubernetes with the NVIDIA GPU Operator).
- Analyze performance issues across compute, network (InfiniBand/RoCE) and storage, and perform root cause analysis.
- Manage cluster operations through NVIDIA Base Command Manager or equivalent tooling.
- Configure and tune monitoring and alerting (DCGM, Prometheus, Grafana) to reduce noise and improve detection.
- Execute changes, patching, firmware and driver updates under change control.
- Mentor L1 engineers and review the quality of escalations.
Skills
- Strong Linux administration and shell scripting (Bash, Python).
- Hands-on experience with GPU diagnostics, DCGM, container runtimes (Docker, containerd) and NVIDIA NGC containers.
- Working knowledge of InfiniBand and NCCL, and of storage for AI (parallel or NFS-based).
L3 NOC Engineer: Senior and Escalation Lead
Experience: 6-10 years in infrastructure or platform operations, including 2+ years on AI, GPU or HPC platforms
Responsibilities
- Act as the final technical escalation point for critical incidents and major outages. Lead incident bridges and post-incident reviews.
- Perform deep diagnostics on multi-node training and inference performance, fabric congestion, NVLink and InfiniBand faults, and GPU memory or ECC errors.
- Design and improve NOC monitoring architecture, observability, alert thresholds and SLO/SLA reporting for AI infrastructure.
- Automate operations using Python, Ansible and Terraform. Build self-healing workflows and reduce manual toil.
- Own problem management, capacity planning and lifecycle management (firmware, driver and CUDA compatibility matrices).
- Engage with NVIDIA and OEM support (Dell, HPE, Supermicro and others), and manage vendor cases through resolution.
- Define runbooks, standards and training plans for L1 and L2, and contribute to client solution reviews and pre-sales technical input where needed.
Skills
- Expert-level understanding of the NVIDIA AI stack: DGX, Base Command, DCGM, NCCL, GPU Operator, NGC, and NVIDIA AI Enterprise.
- Advanced Kubernetes and Slurm operations, and InfiniBand fabric management (UFM).
- Strong scripting and automation, with an observability and incident-command mindset.
Common Qualifications
- Bachelor's degree in Computer Science, IT, Electronics or a related field (or equivalent experience).
- The certifications listed for the level above (validity must be current).
- Willingness to work rotational shifts, weekends and on-call schedules.
- Strong written and verbal communication in English.
Nice to Have
- ITIL 4 Foundation
- Cloud certifications (AWS, Azure or GCP)
- CCNA or CCNP
- RHCSA or CKA
- Experience with air-gapped or on-prem AI environments
- Experience with the ServiceNow, Zabbix, SolarWinds or Nagios toolsets