Network Engineer (AI GPU Cluster Operations)

Aquila Hash, Inc.

Buffalo (NY)

On-site

USD 110,000 - 160,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Aquila Hash, Inc. in Buffalo, NY, is seeking an AI GPU Cluster Operations – Network Engineer to lead operations, maintenance and troubleshooting of high‑performance networking infrastructure for large‑scale AI GPU clusters.

The role centers on InfiniBand, RoCE and data center IP networks, with emphasis on GPU servers, Linux and overall AI cluster architecture. You will monitor health, resolve incidents and work with vendors to drive issues to resolution.

Qualifications

  • Degree or relevant educational background in Computer Science, Information Technology, Networking, Engineering or a related field.
  • 3+ years of hands-on network operations, data center networking or infrastructure experience.
  • Hands-on experience operating InfiniBand and/or RoCE networks in GPU, AI, HPC or large-scale data center environments.
  • Experience with NVIDIA/Mellanox InfiniBand or high-speed Ethernet networking equipment is strongly preferred.
  • Understanding of InfiniBand architecture, Subnet Management, partitions, routing, link states and performance counters.
  • Strong understanding of TCP/IP and data center networking.

Responsibilities

  • Operate and maintain high-performance GPU cluster networks, including InfiniBand and RoCE, and standard data center IP networks.
  • Manage and troubleshoot NVIDIA/Mellanox InfiniBand switches, high-speed Ethernet switches, NICs/HCAs, optical transceivers and related infrastructure.
  • Monitor network health and identify link degradation, congestion, packet loss, bandwidth or latency issues.
  • Perform InfiniBand health checks and diagnostics using ibstat, ibdiagnet and NVIDIA/Mellanox utilities.
  • Support InfiniBand Subnet Management, partition configuration, routing, and performance-counter analysis.
  • Troubleshoot RoCE environments, including connectivity and lossless Ethernet issues.
  • Support data center IP network operations (BGP, OSPF, VLAN, MLAG, TCP/IP).
  • Troubleshoot GPU node connectivity and determine Root Cause across server, NIC/HCA, optics, cabling or fabric.
  • Collaborate with GPU Cluster Operations Engineers to resolve NCCL failures and performance degradation.
  • Monitor via Prometheus, SNMP, Grafana and related platforms; develop monitoring and diagnostic capabilities.
  • Perform approved network configuration changes and SOP-driven maintenance/upgrades.
  • Support network capacity planning and expansion of GPU cluster infrastructure.
  • Automate tasks using Python, Shell, Ansible and/or Terraform.
  • Participate in incident response, post-incident reviews and RCA.
  • Maintain network diagrams, configuration docs and troubleshooting guides.
  • Coordinate with NVIDIA/Mellanox, server OEMs and vendors to resolve hardware/network issues.
  • Support hardware replacement and RMA activities for switches, NICs/HCAs and optics.

Skills

InfiniBand
RoCE
Linux
TCP/IP
Prometheus
Grafana
Python
Shell
Ansible
Terraform
BGP
OSPF
VLAN
MLAG
NCCL

Education

Degree in Computer Science / IT / Networking / Engineering or related field

Tools

Mellanox InfiniBand
ibstat
ibdiagnet
NICs/HCAs
100G/200G/400G NICs

Job description

The AI GPU Cluster Operations – Network Engineer is responsible for the operations, maintenance and troubleshooting of high-performance networking infrastructure supporting large-scale AI GPU computing clusters.

The role focuses on InfiniBand, RoCE and data center IP networks, while also requiring a strong understanding of GPU servers, Linux and the overall AI cluster architecture.

The engineer will monitor network and cluster health, troubleshoot connectivity and performance issues, maintain high-speed network infrastructure, support production incidents, implement approved network changes and work with internal engineering teams and OEM vendors to drive technical issues through resolution.

Key Responsibilities
  • Operate and maintain high-performance GPU cluster networks, including InfiniBand and RoCE, as well as standard data center IP networks.
  • Manage and troubleshoot NVIDIA/Mellanox InfiniBand switches, high-speed Ethernet switches, NICs/HCAs, optical transceivers, cables and related network infrastructure.
  • Monitor network health and identify link degradation, link flapping, congestion, packet loss, bandwidth or latency issues.
  • Perform InfiniBand health checks and diagnostics using tools such as ibstat, ibdiagnet and other NVIDIA/Mellanox diagnostic utilities.
  • Support InfiniBand Subnet Management, partition configuration, routing, link-health monitoring and performance-counter analysis.
  • Troubleshoot RoCE environments, including connectivity, congestion and lossless Ethernet-related issues.
  • Support data center IP network operations, including BGP, OSPF, VLAN, MLAG and TCP/IP.
  • Troubleshoot GPU node connectivity and determine whether infrastructure issues originate from the server, NIC/HCA, optics, cabling, switch or network fabric.
  • Work with GPU Cluster Operations Engineers to troubleshoot GPU-to-GPU and node-to-node communication issues, including network-related NCCL failures and performance degradation.
  • Monitor network and infrastructure health using Prometheus, SNMP, Grafana and related monitoring platforms.
  • Develop and maintain network monitoring, alerting and diagnostic capabilities.
  • Perform approved network configuration changes, maintenance and upgrades according to established change-management procedures and Standard Operating Procedures (SOPs).
  • Support network capacity planning and expansion of GPU cluster infrastructure.
  • Use Python, Shell, Ansible and/or Terraform to automate network configuration, health checks, monitoring and repetitive operational tasks.
  • Participate in infrastructure incident response, troubleshooting, post-incident reviews and Root Cause Analysis (RCA).
  • Maintain network diagrams, configuration documentation, SOPs, troubleshooting guides and operational records.
  • Coordinate with NVIDIA/Mellanox, server OEMs, network vendors and other technical support teams to resolve hardware and network issues.
  • Support hardware replacement and RMA activities for switches, NICs/HCAs, optical transceivers and other network components.
Qualifications
  • Degree or relevant educational background in Computer Science, Information Technology, Networking, Telecommunications, Engineering or a related field.
  • 3+ years of hands-on network operations, data center networking or infrastructure experience.
  • Hands-on experience operating InfiniBand and/or RoCE networks in GPU, AI, HPC or large‑scale data center environments.
  • Experience with NVIDIA/Mellanox InfiniBand or high‑speed Ethernet networking equipment is strongly preferred.
  • Understanding of InfiniBand architecture, including Subnet Management, partitions, routing, link states and performance counters.
  • Strong understanding of TCP/IP and data center networking.
  • Working knowledge of BGP, OSPF, VLAN and MLAG.
  • Ability to troubleshoot high‑speed network connectivity, link degradation, congestion, packet loss and performance issues.
  • Hands‑on ability to troubleshoot switches, NICs/HCAs, optical transceivers, cables and switch ports.
  • Familiarity with InfiniBand diagnostic tools such as ibstat and ibdiagnet.
  • Familiarity with Linux and command‑line troubleshooting.
  • Experience with monitoring platforms and protocols such as Prometheus, SNMP and Grafana.
  • Ability to write basic Python and/or Shell scripts for troubleshooting and automation.
  • Experience with Ansible and/or Terraform is preferred.
  • Good understanding of GPU cluster architecture and the relationship between GPU servers, NICs/HCAs and high‑performance network fabrics.
  • Strong troubleshooting, documentation and Root Cause Analysis (RCA) capabilities.
  • Strong sense of ownership and ability to work effectively during production incidents.
Preferred Qualifications
  • Experience supporting large‑scale 1,000+ GPU clusters; experience with 10,000+ GPU environments is a strong plus.
  • Experience operating 100G/200G/400G or higher‑speed InfiniBand or Ethernet networks.
  • Experience troubleshooting NCCL and GPU communication issues from the network/fabric perspective.
  • Experience with NVIDIA/Mellanox network management and diagnostic tools.
  • Experience with network automation using Python, Shell, Ansible or Terraform.
  • CCNP, CCIE or other relevant networking certifications are a plus.
  • Experience working within ITIL‑based incident, problem and change‑management processes.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Network Engineer
Network Engineer

asobbi • United States

On-site
USD 160,000 - 190,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Network Operations Engineer
Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Network Operations Engineer, AI Networking
Network Operations Engineer, AI Networking

OpenAI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Network Engineer (Supercomputer Infrastructure)
Network Engineer (Supercomputer Infrastructure)

Spacex • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Infra Network Operation Manager
Infra Network Operation Manager

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
AI Operations & Infrastructure Engineer
AI Operations & Infrastructure Engineer

Invictus International Consulting, LLC • Fort Meade (MD)

On-site
USD 100,000 - 130,000