Senior AI Data Center Fabric Architect

NVIDIA Corporation

England

On-site

GBP 90,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA Corporation is seeking a Senior Network and Fabric Operations Solutions Architect to join its Infrastructure Specialist Team. You will own Day 2 fabric operations, manage NVLink/NVSwitch partitions on large GPU estates, and lead network fabric validation and multi-tenant workflows, engaging with customers and partners.

The role requires 5+ years in data center networking, deep understanding of InfiniBand, RoCE, Cumulus Linux/SONiC, and experience with Python/Bash, IaC tools, and

Qualifications

  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
  • 5+ years of professional experience in data centre networking, fabric engineering or large-scale network operations roles.
  • Deep understanding of data centre network architectures and RDMA fabrics—InfiniBand and RoCE/Ethernet.
  • Hands-on experience with InfiniBand and UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX and BlueField NICs/DPUs, and NVLink/NVSwitch systems.
  • Linux and switch OS internals, host networking and OS-level security.
  • Python and Bash scripting, Infra as Code tools (Ansible, Terraform), GitOps, and observability stacks (Grafana, Prometheus).
  • Experience delivering multi-tenant network isolation for GPU estates and performing architectural reviews for executive stakeholders.

Responsibilities

  • Own Day 2 fabric operations across NVIDIA GPU fleets and partition management.
  • Own the switch software and firmware lifecycle with upgrade plans across partner fabrics.
  • Support Day 1 fabric validation including InfiniBand bring-up and network configuration.
  • Minimise time from cluster handover to first production workload and ensure reliability at fleet scale.
  • Provide consultative guidance and hands-on troubleshooting across the fabric stack and runbooks for partner teams.

Skills

Networking expertise
Consulting skills
Kubernetes networking
Python scripting
Bash scripting

Education

BS/MS/PhD in Computer Science or related

Tools

Cumulus Linux
SONiC
InfiniBand
UFM
Spectrum-X
NVLink/NVSwitch

Job description

NVIDIA Corporation is seeking a Senior Network and Fabric Operations Solutions Architect to join its Infrastructure Specialist Team. You will own Day 2 fabric operations, manage NVLink/NVSwitch partitions on large GPU estates, and lead network fabric validation and multi-tenant workflows, engaging with customers and partners.

The role requires 5+ years in data center networking, deep understanding of InfiniBand, RoCE, Cumulus Linux/SONiC, and experience with Python/Bash, IaC tools, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network Fabric Architect for AI Data Centers
Senior Network Fabric Architect for AI Data Centers

NVIDIA Corporation • Reading

On-site
GBP 120,000 - 180,000
Senior Cloud Fabric Architect - AI/DL GPU Networking
Senior Cloud Fabric Architect - AI/DL GPU Networking

NVIDIA • Reading

On-site
GBP 110,000 - 150,000
Senior Cloud Infrastructure and Network Operations Solutions Architect
Senior Cloud Infrastructure and Network Operations Solutions Architect

NVIDIA • Reading

On-site
GBP 110,000 - 150,000
Senior Cloud Infrastructure and Network Operations Solutions Architect
Senior Cloud Infrastructure and Network Operations Solutions Architect

NVIDIA Corporation • Reading

On-site
GBP 120,000 - 180,000
Staff Network Architect – AI/HPC Fabric & Automation
Staff Network Architect – AI/HPC Fabric & Automation

Nscale • Greater London

On-site
GBP 90,000 - 130,000
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC

NVIDIA • United Kingdom

On-site
GBP 110,000 - 170,000
HPC-AI Cluster Architect & Automation Lead
HPC-AI Cluster Architect & Automation Lead

NVIDIA • United Kingdom

On-site
GBP 90,000 - 120,000
HPC Network Engineer: RDMA, Leaf-Spine & AI Fabric
HPC Network Engineer: RDMA, Leaf-Spine & AI Fabric

Fuse Energy • Greater London

On-site
GBP 90,000 - 140,000
Biannual bonus
Fully expensed tech
Private health insurance
+1
Senior HPC Network Architect for GPU-Driven AI Cloud
Senior HPC Network Architect for GPU-Driven AI Cloud

lambda • Manchester

Hybrid
GBP 106,000 - 159,000
Cash & equity compensation
Health, dental, vision coverage
Wellness and commuter stipends
+1
Network Architect (GPU AI Data Centres)
Network Architect (GPU AI Data Centres)

Experis UK • Belfast City District

Hybrid
GBP 90,000 - 130,000
Hybrid work model
On-site facilities