Senior Network Engineer – GPU Cluster Networking

Advanced Micro Devices

San Jose (CA)

Hybrid

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Advanced Micro Devices (AMD) seeks a Senior Network Engineer to architect and operate high‑performance backend networks for large GPU clusters, including 10,000+ GPUs. You will own the fabric from GPU servers to leaf-spine, optimize RoCEv2 Ethernet fabrics, and model capacity for AI/HPC workloads.

You will collaborate with AI, data center, storage, and security teams, lead incident response, and mentor engineers while driving scalable, observable networking solutions.

Qualifications

  • Significant experience designing, deploying, and operating production data center networks for AI/GPU/HPC environments.
  • Experience building backend networks for GPU clusters containing ~10,000+ GPUs or similar hyperscale systems.
  • Deep knowledge of data center networking fundamentals and modern leaf-spine architectures.
  • Hands-on experience with RDMA, RoCEv2, PFC, DCQCN, QoS, and related technologies.

Responsibilities

  • Architect, deploy, and operate high-performance back-end networks for large GPU clusters.
  • Own end-to-end network path from GPU servers/NICs through leaf-spine fabric.
  • Design RoCEv2 Ethernet fabrics and scalable topologies; model capacity and plan growth.
  • Lead incident response, root-cause analysis, and optimization for network workloads.
  • Collaborate across AI, data center, storage, and security teams to prevent bottlenecks.

Skills

RDMA
RoCEv2
Data center networking
Leaf-spine
GPU clusters
BGP/ECMP
EVPN/VXLAN
Prometheus
Grafana
Junos OS
NVIDIA/AMD ROCm RCCL

Education

Bachelor's or Master's in Computer Engineering

Tools

Juniper Junos OS
Prometheus
Grafana
RDMA tooling

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou'redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger- technologythat moves the world forward.Join us and, together, we'll advance your career.

THE ROLE:

We are seeking a Senior Network Engineer to join the AMD IT System Engineering team.

This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.

The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.

The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.

You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.

THE PERSON:

You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.

You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.

You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.

KEY RESPONSIBILITIES
  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
  • Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
  • Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
  • Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
  • Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
    • Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
  • Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
  • Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
  • Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.
  • Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations
PREFERRED EXPERIENCE:
  • Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
  • Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
  • Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments.
  • Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
  • Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
  • Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
  • Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
  • Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.
  • Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.
  • Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.
  • Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.
ACADEMIC CREDITALS:
  • Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience.
LOCATION:

San Jose, CA OR Austin, TX

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer – GPU Cluster Networking
Senior Network Engineer – GPU Cluster Networking

AMD • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Hybrid work model
AMD benefits
Data Center Infrastructure Architect
Data Center Infrastructure Architect

CareerArc • North Carolina

On-site
USD 150,000 - 230,000
Product Manager - Network Infrastructure
Product Manager - Network Infrastructure

Advanced Micro Devices, Inc. • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Hybrid work environment
Product Manager - Network Infrastructure
Product Manager - Network Infrastructure

AMD • Santa Clara (CA)

Hybrid
USD 130,000 - 160,000
Competitive salary
Hybrid work model
Comprehensive benefits
Product Application Engineer- Data Center Deployment
Product Application Engineer- Data Center Deployment

AMD • California (MO)

On-site
USD 80,000 - 110,000
Health insurance
401(k) plan
Paid time off
Cloud & Customer Solutions Engineer - DC GPU
Cloud & Customer Solutions Engineer - DC GPU

Advanced Micro Devices • Bellevue (WA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Sr. Staff Software Development Engineer - Collectives and Network optimization
Sr. Staff Software Development Engineer - Collectives and Network optimization

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 130,000 - 160,000
Comprehensive benefits package
Technical Marketing Engineer – Datacenter GPU Clusters & Networking
Technical Marketing Engineer – Datacenter GPU Clusters & Networking

AMD • Santa Clara (CA)

On-site
USD 140,000 - 190,000
AMD Benefits
Senior AI Switch Systems Design Engineer
Senior AI Switch Systems Design Engineer

Advanced Micro Devices • Secaucus (NJ)

Hybrid
USD 100,000 - 130,000
Inclusive workplace
AMD benefits package
AI/HPC Cluster Design Engineer
AI/HPC Cluster Design Engineer

Socket.dev • Austin (TX)

On-site
USD 140,000 - 230,000