Senior Network Architect for 10k+ GPU HPC Clusters

AMD

San Jose (CA)

Hybrid

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
AMD benefits

Job summary

AMD is seeking a Senior Network Engineer in San Jose, CA to architect, deploy, and optimize backend networks for expansive GPU clusters. You will own the network from GPU servers to the leaf-spine fabric, ensuring low latency and high bandwidth for AI/HPC workloads.

The role involves design, qualification, and production deployment with cross-team collaboration. You will lead complex initiatives, mentor engineers, and build observability with Prometheus/Grafana, while working in a hybrid

Qualifications

  • Bachelor’s degree in Computer Engineering or related field.
  • Extensive experience in data center networking for AI/HPC workloads.
  • Deep knowledge of RDMA, RoCEv2 and large-scale GPU clusters.

Responsibilities

  • Architect and operate high-performance backend networks for large AMD Instinct GPU clusters.
  • Design fabrics supporting AI/HPC from rack to 10,000+ GPUs.
  • Own backend network from GPU server to leaf-spine fabric.
  • Develop RoCEv2-based Ethernet fabrics and high-speed topology.
  • Lead incident response and capacity planning for GPU networks.
  • Collaborate with AI engineering, data center, storage, and security teams.

Skills

Data center networking
RDMA/RoCEv2
GPU cluster design
Telemetry & performance testing
Leadership & mentoring

Education

Bachelor’s or Master’s in Computer Engineering or related field

Tools

Juniper Junos OS
Prometheus
Grafana
VXLAN/EVPN
Kubernetes

Job description

AMD is seeking a Senior Network Engineer in San Jose, CA to architect, deploy, and optimize backend networks for expansive GPU clusters. You will own the network from GPU servers to the leaf-spine fabric, ensuring low latency and high bandwidth for AI/HPC workloads.

The role involves design, qualification, and production deployment with cross-team collaboration. You will lead complex initiatives, mentor engineers, and build observability with Prometheus/Grafana, while working in a hybrid

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Architect – HPC GPU Data Center
Senior Network Architect – HPC GPU Data Center

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Senior Network Engineer – GPU Cluster Networking
Senior Network Engineer – GPU Cluster Networking

AMD • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Hybrid work model
AMD benefits
Senior Network Engineer – GPU Cluster Networking
Senior Network Engineer – GPU Cluster Networking

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Senior AI Network Architect for Large-Scale GPU Clusters
Senior AI Network Architect for Large-Scale GPU Clusters

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Senior GPU Compute Infra Architect - Onsite SF
Senior GPU Compute Infra Architect - Onsite SF

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
Senior HPC & Infiniband Network Engineer
Senior HPC & Infiniband Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior HPC Architect: At-Scale GPU Deployment & Automation
Senior HPC Architect: At-Scale GPU Deployment & Automation

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior GPU HPC Systems Engineer
Senior GPU HPC Systems Engineer

Acceler8 Talent • Fremont (CA)

On-site
USD 135,000 - 165,000
Comprehensive benefits
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits