Senior Network Engineer - GPU Cluster Networking

Advanced Micro Devices (AMD)

Hyderabad

On-site

INR 2,500,000 - 4,000,000

Full time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Advanced Micro Devices (AMD) is seeking a Senior Network Engineer to join the IT System Engineering team in Hyderabad. You will architect, deploy, and optimize high-performance backend networks for AMD Instinct GPU clusters, ensuring scalable bandwidth and low latency for AI training workloads.

The role emphasizes RoCEv2, leaf-spine fabric, and modern data center technologies, with ownership of end-to-end network performance, incident response, capacity planning, and improvement across globally

Qualifications

  • Hands-on data center networking and GPU cluster experience.
  • RDMA/RoCEv2 expertise in production environments.
  • Strong understanding of network topology, leaf-spine, and EVPN/VXLAN.

Responsibilities

  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD GPU clusters.
  • Own the backend network architecture from GPU servers through the leaf-spine fabric.
  • Configure and optimize RoCEv2, PFC, ECN, QoS for lossless or near-lossless Ethernet.
  • Lead production incident response, root-cause analysis, and capacity planning.
  • Document architecture and operating standards across globally distributed teams.

Education

Bachelor's or Master's in Computer Engineering

Tools

Juniper
Prometheus
Grafana
gNMI
SNMP

Job description

Senior Network Engineer - GPU Cluster Networking
The Role

We are seeking a Senior Network Engineer to join the AMD IT System Engineering team. This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.

The Person

You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs.

You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.

You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.

Key Responsibilities
  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
  • Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
  • Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
  • Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
  • Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
  • Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
  • Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
  • Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
  • Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.
  • Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations
Preferred Experience
  • Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
  • Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
  • Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments.
  • Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
  • Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
  • Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
  • Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
  • Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.
  • Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.
  • Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.
  • Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.
Academic Credentials
  • Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience.
Location

Location: Hyderabad

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network Engineer – GPU Cluster Networking
Senior Network Engineer – GPU Cluster Networking

AMD • Hyderabad

On-site
INR 3,500,000 - 6,200,000
Senior Network Engineer – GPU Cluster Networking
Senior Network Engineer – GPU Cluster Networking

Advanced Micro Devices • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Senior Data Center Network Engineer (Fabric & RoCE) | Confidential Technology Company | Noida, India
Senior Data Center Network Engineer (Fabric & RoCE) | Confidential Technology Company | Noida, India

Confidential Technology Company • Dadri

On-site
INR 3,500,000 - 5,600,000
Senior Solutions Architect Networking And Compute Infrastructure
Senior Solutions Architect Networking And Compute Infrastructure

NVIDIA AI • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA • Gurugram District

On-site
INR 3,500,000 - 7,500,000
Senior Network Engineer
Senior Network Engineer

Tenarai • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Senior Solutions Architect, Networking and Compute Infrastructure
Senior Solutions Architect, Networking and Compute Infrastructure

NVIDIA Gruppe • Gurugram District

On-site
INR 3,000,000 - 6,000,000
Product Application Engineer - Networking
Product Application Engineer - Networking

Advanced Micro Devices (AMD) • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Senior Software Engineer Fabric Networking - GPU
Senior Software Engineer Fabric Networking - GPU

NVIDIA • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Competitive compensation
Industry-leading benefits
Collaborative culture
+1
Senior Network Engineer
Senior Network Engineer

Operations • Mumbai

On-site
INR 600,000 - 1,200,000