The AI GPU Cluster Operations – Network Engineer is responsible for the operations, maintenance and troubleshooting of high-performance networking infrastructure supporting large-scale AI GPU computing clusters.
The role focuses on InfiniBand, RoCE and data center IP networks, while also requiring a strong understanding of GPU servers, Linux and the overall AI cluster architecture.
The engineer will monitor network and cluster health, troubleshoot connectivity and performance issues, maintain high-speed network infrastructure, support production incidents, implement approved network changes and work with internal engineering teams and OEM vendors to drive technical issues through resolution.
Key Responsibilities
- Operate and maintain high-performance GPU cluster networks, including InfiniBand and RoCE, as well as standard data center IP networks.
- Manage and troubleshoot NVIDIA/Mellanox InfiniBand switches, high-speed Ethernet switches, NICs/HCAs, optical transceivers, cables and related network infrastructure.
- Monitor network health and identify link degradation, link flapping, congestion, packet loss, bandwidth or latency issues.
- Perform InfiniBand health checks and diagnostics using tools such as ibstat, ibdiagnet and other NVIDIA/Mellanox diagnostic utilities.
- Support InfiniBand Subnet Management, partition configuration, routing, link-health monitoring and performance-counter analysis.
- Troubleshoot RoCE environments, including connectivity, congestion and lossless Ethernet-related issues.
- Support data center IP network operations, including BGP, OSPF, VLAN, MLAG and TCP/IP.
- Troubleshoot GPU node connectivity and determine whether infrastructure issues originate from the server, NIC/HCA, optics, cabling, switch or network fabric.
- Work with GPU Cluster Operations Engineers to troubleshoot GPU-to-GPU and node-to-node communication issues, including network-related NCCL failures and performance degradation.
- Monitor network and infrastructure health using Prometheus, SNMP, Grafana and related monitoring platforms.
- Develop and maintain network monitoring, alerting and diagnostic capabilities.
- Perform approved network configuration changes, maintenance and upgrades according to established change-management procedures and Standard Operating Procedures (SOPs).
- Support network capacity planning and expansion of GPU cluster infrastructure.
- Use Python, Shell, Ansible and/or Terraform to automate network configuration, health checks, monitoring and repetitive operational tasks.
- Participate in infrastructure incident response, troubleshooting, post-incident reviews and Root Cause Analysis (RCA).
- Maintain network diagrams, configuration documentation, SOPs, troubleshooting guides and operational records.
- Coordinate with NVIDIA/Mellanox, server OEMs, network vendors and other technical support teams to resolve hardware and network issues.
- Support hardware replacement and RMA activities for switches, NICs/HCAs, optical transceivers and other network components.
Qualifications
- Degree or relevant educational background in Computer Science, Information Technology, Networking, Telecommunications, Engineering or a related field.
- 3+ years of hands-on network operations, data center networking or infrastructure experience.
- Hands-on experience operating InfiniBand and/or RoCE networks in GPU, AI, HPC or large‑scale data center environments.
- Experience with NVIDIA/Mellanox InfiniBand or high‑speed Ethernet networking equipment is strongly preferred.
- Understanding of InfiniBand architecture, including Subnet Management, partitions, routing, link states and performance counters.
- Strong understanding of TCP/IP and data center networking.
- Working knowledge of BGP, OSPF, VLAN and MLAG.
- Ability to troubleshoot high‑speed network connectivity, link degradation, congestion, packet loss and performance issues.
- Hands‑on ability to troubleshoot switches, NICs/HCAs, optical transceivers, cables and switch ports.
- Familiarity with InfiniBand diagnostic tools such as ibstat and ibdiagnet.
- Familiarity with Linux and command‑line troubleshooting.
- Experience with monitoring platforms and protocols such as Prometheus, SNMP and Grafana.
- Ability to write basic Python and/or Shell scripts for troubleshooting and automation.
- Experience with Ansible and/or Terraform is preferred.
- Good understanding of GPU cluster architecture and the relationship between GPU servers, NICs/HCAs and high‑performance network fabrics.
- Strong troubleshooting, documentation and Root Cause Analysis (RCA) capabilities.
- Strong sense of ownership and ability to work effectively during production incidents.
Preferred Qualifications
- Experience supporting large‑scale 1,000+ GPU clusters; experience with 10,000+ GPU environments is a strong plus.
- Experience operating 100G/200G/400G or higher‑speed InfiniBand or Ethernet networks.
- Experience troubleshooting NCCL and GPU communication issues from the network/fabric perspective.
- Experience with NVIDIA/Mellanox network management and diagnostic tools.
- Experience with network automation using Python, Shell, Ansible or Terraform.
- CCNP, CCIE or other relevant networking certifications are a plus.
- Experience working within ITIL‑based incident, problem and change‑management processes.