Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia

Sassari

On-site

EUR 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA is seeking a Senior Cloud Infrastructure and Network Operations Architect to own Day 2 fabric operations for Cloud Partner estates, guiding InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch across multi-tenant AI/HPC systems.

You will lead hardware bring-up validation, fleet reliability, and runbooks for partner teams, with hands-on tuning of NICs/DPUs, switch OS, and Kubernetes networking in a fast, open-source–first fabric ecosystem.

Qualifications

  • 5+ years in data center networking and large-scale operations.
  • Deep knowledge of InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch.
  • Experience with NVL72-class platforms and NICs/DPUs.
  • Strong Linux (RedHat/Ubuntu) and switch OS skills.
  • Automation skills: Python, Bash, Ansible, Terraform.
  • Kubernetes networking in GPU clusters and multi-tenant isolation.

Responsibilities

  • Own Day 2 fabric operations across NVIDIA Cloud Partner fleets (NVLink/NVSwitch partition management, maintenance-partition isolation, safe partition-change workflows)
  • Manage switch software and firmware lifecycle including upgrades and rollout campaigns with minimal production disruption
  • Support Day 1 fabric validation and acceptance (InfiniBand/UFM bring-up, Spectrum-X/RoCE config, burn-in against MTBI/goodput targets)
  • Minimise handover time to production by reducing duplicated validation across hardware bring-up and partner operations
  • Drive fleet-wide fabric reliability through telemetry, fault detection, remediation, and root-cause analysis of network-induced job failures
  • Provide consultative guidance and hands-on troubleshooting across fabric stack (NICs/DPUs, switch OS, routing, Kubernetes networking) and lead knowledge transfer with runbooks for partner teams
  • Serve as technical leader for assigned accounts and present to executive stakeholders
  • Deep understanding of data center architectures and RDMA fabrics
  • Hands-on experience with InfiniBand/UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC
  • Strong Linux knowledge and HPC/AI traffic patterns
  • Automation and Observability skills: Python/Bash, IaC tools (Ansible, Terraform), GitOps, Grafana/Loki/Prometheus
  • Proven ability to measure and improve MTBI and job goodput, and lead architectural reviews with executives
  • Experience with Kubernetes networking in GPU clusters and multi-tenant network isolation concepts
  • Strong consultative mindset with knowledge transfer and runbook creation
  • Strong communication and presentation skills
  • Cross-functional collaboration
  • InfiniBand and UFM
  • Spectrum-X Ethernet

Skills

Data center networking
Fabric engineering
NVLink/NVSwitch
InfiniBand
RoCE/Ethernet
RDMA fabrics
Linux
Automation
Python
Ansible
Terraform
Grafana
Prometheus
Kubernetes networking
Multi-tenant isolation
UFM
Spectrum-X
Cumulus Linux
SONiC
NICs/DPUs
ConnectX/BlueField
NVL72

Tools

Cumulus Linux
SONiC
ConnectX/BlueField
NICs/DPUs
Kubernetes

Job description

In this role you will own Day 2 fabric operations for NVIDIA’s Cloud Partner estates, guiding the reliability and performance of InfiniBand, RoCE/Ethernet, and NVLink/NVSwitch fabrics at scale. You’ll work with cross-functional teams and customers to architect, validate, and operate complex GPU-centric networks across multi-tenant environments. This is a hands-on leadership role in a fast-moving, open-source–first fabric ecosystem. You’ll drive operational excellence, incident response, and knowledge transfer to partner teams, shaping how large AI/HPC systems run reliably.

  • Own Day 2 fabric operations across NVIDIA Cloud Partner fleets (NVLink/NVSwitch partition management, maintenance-partition isolation, safe partition-change workflows)
  • Manage switch software and firmware lifecycle including upgrades and rollout campaigns with minimal production disruption
  • Support Day 1 fabric validation and acceptance (InfiniBand/UFM bring-up, Spectrum-X/RoCE config, burn-in against MTBI/goodput targets)
  • Minimise handover time to production by reducing duplicated validation across hardware bring-up and partner operations
  • Drive fleet-wide fabric reliability through telemetry, fault detection, remediation, and root-cause analysis of network-induced job failures
  • Provide consultative guidance and hands-on troubleshooting across fabric stack (NICs/DPUs, switch OS, routing, Kubernetes networking) and lead knowledge transfer with runbooks for partner teams
  • Serve as technical leader for assigned accounts and present to executive stakeholders
  • 5+ years in data center networking, fabric engineering, or large-scale network operations
  • Deep understanding of data center architectures and RDMA fabrics (InfiniBand, RoCE/Ethernet) including topology, routing, congestion control, and QoS
  • Hands-on experience with InfiniBand/UFM, Spectrum-X Ethernet, Cumulus Linux and/or SONiC, ConnectX/BlueField NICs/DPUs, NVLink/NVSwitch on NVL72-class platforms
  • Strong Linux knowledge (RedHat/Ubuntu), switch OS internals, security, and HPC/AI traffic patterns
  • Automation and Observability skills: Python/Bash, IaC tools (Ansible, Terraform), GitOps, Grafana/Loki/Prometheus
  • Proven ability to measure and improve MTBI and job goodput, and to lead architectural reviews with executive stakeholders
  • Experience with Kubernetes networking in GPU clusters and multi-tenant network isolation concepts
  • Strong consultative mindset with demonstrated capacity to transfer knowledge and create runbooks for partner teams
  • Consultative mindset
  • Strong communication and presentation skills
  • Cross-functional collaboration
  • InfiniBand and UFM
  • Spectrum-X Ethernet

Senior Cloud Infrastructure and Network Operations Solutions Architect sassari, sardegna, it

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Caserta

On-site
EUR 120,000 - 180,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Macerata

On-site
EUR 90,000 - 130,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Firenze

On-site
EUR 120,000 - 180,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Rovigo

On-site
EUR 110,000 - 150,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Udine

On-site
EUR 90,000 - 150,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Perugia

On-site
EUR 90,000 - 130,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Piemonte

On-site
EUR 120,000 - 160,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Teramo

On-site
EUR 120,000 - 180,000
Senior Cloud Infrastructure And Network Operations Solutions Architect
Senior Cloud Infrastructure And Network Operations Solutions Architect

Nvidia • Rimini

On-site
EUR 140,000 - 200,000
Senior Cloud Fabric Operations Architect for GPU AI
Senior Cloud Fabric Operations Architect for GPU AI

Nvidia • Rimini

On-site
EUR 140,000 - 200,000