Senior Cloud Fabric Architect - AI/DL GPU Networking

NVIDIA

Reading

On-site

GBP 110,000 - 150,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

NVIDIA is seeking a Senior Network and Fabric Operations Solutions Architect to join its Infrastructure Specialist Team. You will own Day 2 fabric operations across large GPU fleets, guiding both design and hands-on deployment.

The role combines deep networking, consulting, and leadership, engaging with customers and partners to architect, validate, and operate InfiniBand, Ethernet, NVLink/NVSwitch and related fabrics in diverse data-centre environments.

Qualifications

  • B.S./M.S./PhD or equivalent in a related field or experience.
  • 5+ years in data centre networking, fabric engineering or large-scale network operations.
  • Deep understanding of InfiniBand/RDMA fabrics and RoCE/Ethernet.
  • Hands-on NVIDIA fabric tech: InfiniBand, UFM, Spectrum-X, Cumulus/SONiC, NVLink/NVSwitch.
  • Strong Linux knowledge and switch OS concepts; security and HPC/AI traffic patterns.
  • Proficiency in Python, Bash, IaC (Ansible/Terraform), and observability stacks.

Responsibilities

  • Own Day 2 fabric operations across NCP fleets, including NVLink/NVSwitch partitioning and multi-tenant isolation.
  • Own switch software lifecycle: Cumulus Linux, SONiC, upgrades, and staged rollout campaigns.
  • Support Day 1 validation: InfiniBand bring-up, RoCE config, and burn-in against MTBI targets.
  • Minimise handover-to-production time by unifying validation across bring-up, managed services, and partner ops.
  • Drive fabric reliability at fleet scale: telemetry, fault detection, remediation, and MTBI improvement.
  • Provide consultative guidance across NICs, DPUs, switch OS, Kubernetes networking, and runbooks for partner teams.

Skills

Fabric & Data Centre Networking
Linux & Switch Platforms
Automation, GitOps & Observability
Fleet Reliability & Customer Engagment
Kubernetes networking in GPU clusters

Education

BS/MS/PhD in Computer Science or related fields

Tools

InfiniBand/UFM
Spectrum-X Ethernet
Cumulus Linux/SONiC
NVLink/NVSwitch
DOCA/DPUs
Grafana/Prometheus/Loki

Job description

NVIDIA is seeking a Senior Network and Fabric Operations Solutions Architect to join its Infrastructure Specialist Team. You will own Day 2 fabric operations across large GPU fleets, guiding both design and hands-on deployment.

The role combines deep networking, consulting, and leadership, engaging with customers and partners to architect, validate, and operate InfiniBand, Ethernet, NVLink/NVSwitch and related fabrics in diverse data-centre environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Data Center Fabric Architect
Senior AI Data Center Fabric Architect

NVIDIA Corporation • England

On-site
GBP 90,000 - 160,000
Senior Network Fabric Architect for AI Data Centers
Senior Network Fabric Architect for AI Data Centers

NVIDIA Corporation • Reading

On-site
GBP 120,000 - 180,000
Senior Cloud Infrastructure and Network Operations Solutions Architect
Senior Cloud Infrastructure and Network Operations Solutions Architect

NVIDIA • Reading

On-site
GBP 110,000 - 150,000
Senior Cloud Infrastructure and Network Operations Solutions Architect
Senior Cloud Infrastructure and Network Operations Solutions Architect

NVIDIA Corporation • Reading

On-site
GBP 120,000 - 180,000
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC
Senior Cloud & DevOps Architect — GPU-Accelerated AI/HPC

NVIDIA • United Kingdom

On-site
GBP 110,000 - 170,000
Staff Network Architect – AI/HPC Fabric & Automation
Staff Network Architect – AI/HPC Fabric & Automation

Nscale • Greater London

On-site
GBP 90,000 - 130,000
Senior HPC Network Architect for GPU-Driven AI Cloud
Senior HPC Network Architect for GPU-Driven AI Cloud

lambda • Manchester

Hybrid
GBP 106,000 - 159,000
Cash & equity compensation
Health, dental, vision coverage
Wellness and commuter stipends
+1
Network Architect (GPU AI Data Centres)
Network Architect (GPU AI Data Centres)

Experis UK • Belfast City District

Hybrid
GBP 90,000 - 130,000
Hybrid work model
On-site facilities
HPC Network Engineer: RDMA, Leaf-Spine & AI Fabric
HPC Network Engineer: RDMA, Leaf-Spine & AI Fabric

Fuse Energy • Greater London

On-site
GBP 90,000 - 140,000
Biannual bonus
Fully expensed tech
Private health insurance
+1
Senior GPU Networking Architect: AI-Scale Kernel Innovator
Senior GPU Networking Architect: AI-Scale Kernel Innovator

NVIDIA • United Kingdom

On-site
GBP 90,000 - 140,000
Competitive salary
Comprehensive benefits package