Network Solutions Architect, AI Factory Services

Lenovo

Morrisville

Hybrid

CAD 199,000 - 271,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Lenovo seeks a highly experienced Network Solutions Architect to join the Hybrid Cloud Solutions and AI Offering Engineering team within SSG. This senior-level role designs, deploys, validates, and optimizes high-performance network infrastructure for Lenovo's AI Factory and GigaFactory services, supporting enterprise-scale GPU environments powered by NVIDIA Spectrum-X Ethernet and InfiniBand fabrics.

The architect will develop reference architectures, runbooks, and field-ready documentation

Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, Information Technology, or related discipline.
  • 5+ years designing and deploying high-performance networking solutions for AI infrastructure.
  • Experience with customer-facing technical consulting, architecture reviews, or deployment leadership.

Responsibilities

  • Design GPU cluster network architectures for NVIDIA AI Factory environments using InfiniBand and Spectrum Ethernet.
  • Lead network bring-up, validation, and production-readiness activities for GPU infrastructure.
  • Develop RDMA/RoCEv2 configuration guides and deployment runbooks for AI and HPC environments.
  • Provide technical leadership during customer engagements, infrastructure deployments, escalations, and architecture reviews.
  • Support global deployments with travel up to 40-50% for workshops and validation.

Skills

Network design
AI infrastructure
Customer-facing consulting

Education

Bachelor's or Master's in CS/EE/IT

Tools

Python
Ansible
Terraform
REST APIs

Job description

* United States of America - North Carolina - Morrisville

Why Work at Lenovo

We are Lenovo. We do what we say. We own what we do. We WOW our customers.

Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).

This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com , and read about the latest news via our StoryHub .

Description and Requirements

Job Summary

Lenovo seeks a highly experienced Network Solutions Architect to join the Hybrid Cloud Solutions and AI Offering Engineering team within SSG. This senior-level role designs, deploys, validates, and optimizes high-performance network infrastructure for Lenovo's AI Factory and GigaFactory service offerings, supporting both enterprise-scale and hyperscale GPU environments powered by NVIDIA Spectrum-X Ethernet and InfiniBand fabrics.

The architect will develop network reference architectures, deployment runbooks, performance validation procedures, and field-ready engineering documentation consumed by Lenovo Professional Services and Managed Services teams globally. This includes GPU cluster fabric design, BlueField DPU architectures, multi-tenant network isolation, high-performance RDMA and RoCEv2 deployments, and production-scale AI infrastructure supporting large distributed training and inference workloads.

The ideal candidate brings deep expertise in NVIDIA Spectrum-X, Quantum InfiniBand, and BlueField DPU technologies, along with hands‑on experience architecting, deploying, validating, and troubleshooting large‑scale GPU clusters. Experience supporting hyperscale AI infrastructure, high‑density liquid‑cooled environments, GPU cluster bring‑up, infrastructure validation, performance tuning, and customer‑facing technical engagements is highly desired.

Key Responsibilities

AI Fabric Architecture & Design

  • Design GPU cluster network architectures for NVIDIA AI Factory environments utilizing:
  • Quantum XDR InfiniBand and Spectrum Ethernet for rack‑scale GigaFactory deployments
  • Design and validate large‑scale AI fabrics supporting RDMA, RoCEv2, GPUDirect, NCCL collectives, and high‑bandwidth GPU‑to‑GPU communications.
  • Develop rail‑optimized InfiniBand topologies and Spectrum‑X Adaptive Routing, SHARP, congestion management, and performance optimization strategies for AI training and inference environments.
  • Contribute to architecture decisions supporting large‑scale distributed AI and HPC workloads.

Deployment, Validation & Performance Engineering

  • Lead network bring‑up, validation, and production‑readiness activities for GPU infrastructure.
  • Develop and validate RDMA/RoCEv2 configuration guides and deployment runbooks for AI and HPC environments.
  • RDMA throughput
  • Storage and network throughput
  • End‑to‑end infrastructure readiness
  • Design failure‑domain isolation strategies and resilient network architectures for large‑scale AI deployments.
  • Support network architecture for high‑density liquid‑cooled GPU environments with power‑aware design considerations.

Troubleshooting & Operational Engineering

  • Diagnose and resolve complex InfiniBand, Ethernet, RDMA, and GPU workload performance issues.
  • Analyze congestion, telemetry, traffic patterns, link utilization, routing behavior, and fabric health to identify bottlenecks and optimize performance.
  • Support root cause analysis and remediation of network, storage, and infrastructure issues impacting AI workload performance.
  • Develop operational runbooks and troubleshooting procedures consumed directly by Lenovo field teams and customers.
  • Design multi‑tenant networking architectures supporting AI Factory, NeoCloud, and managed‑service provider environments.
  • Implement namespace isolation, east‑west traffic segmentation, secure tenant separation, and site resiliency aligned with NVIDIA Cloud Partner Reference Architectures.
  • Design BlueField‑3 DPU solutions utilizing DOCA for infrastructure offload, security services, observability, and service mesh capabilities.

Customer & Cross‑Functional Engagement

  • Collaborate with Lenovo engineering, product, ISG, and services organizations to validate designs against NVIDIA reference architectures and future hardware roadmaps.
  • Work directly with customers, partners, and delivery teams to translate AI workload requirements into production‑ready infrastructure solutions.
  • Provide technical leadership during customer engagements, infrastructure deployments, escalations, and architecture reviews.
  • Support global deployments through approximately 40-50% travel, including customer workshops, implementation support, solution validation, and executive‑level technical discussions.

Basic Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, Information Technology, or related discipline.
  • 5+ years of experience designing and deploying high‑performance networking solutions for:
  • AI infrastructure
  • Experience designing and validating large‑scale network fabrics supporting AI and distributed compute workloads.
  • Experience with customer‑facing technical consulting, architecture reviews, or deployment leadership.

Preferred Qualifications

NVIDIA AI Infrastructure Expertise

  • Expert‑level knowledge of:
  • NVIDIA Spectrum‑X
  • Spectrum‑4 SN5600 Ethernet
  • Quantum InfiniBand (HDR/NDR/XDR)
  • BlueField DPUs
  • DOCA SDK
  • RDMA and RoCEv2

AI Fabric & GPU Cluster Experience

  • Experience supporting large‑scale GPU cluster deployments and production AI environments.
  • Deep understanding of:
  • GPUDirect
  • NCCL
  • GPU cluster validation and performance tuning
  • Adaptive Routing
  • SHARP
  • Congestion control
  • Rail‑optimized InfiniBand architectures

Networking & Infrastructure

  • Strong experience with:
  • BGP
  • EVPN
  • VXLAN
  • MPLS
  • OSPF
  • IS‑IS
  • Experience designing lossless Ethernet environments utilizing:
  • PFC
  • ECN
  • DCQCN
  • QoS
  • Knowledge of high‑performance storage networking and end‑to‑end infrastructure optimization.

Automation & Observability

  • Experience with:
  • Python
  • Ansible
  • Terraform
  • REST APIs
  • Experience implementing infrastructure observability, telemetry, monitoring, and troubleshooting solutions for large‑scale networks.
  • Familiarity with data center operations, performance analytics, and capacity planning.

Certifications

Preferred certifications include:

  • NVIDIA Certified Networking Professional
  • NVIDIA InfiniBand Specialist
  • OCI Networking Specialist
  • AWS Solutions Architect Associate

Hybrid Schedule on campus in Morrisville, NC. 3 days in office, 2 days work from home.

We are an Equal Opportunity Employer and do not discriminate against any employee or applicant for employment because of race, color, sex, age, religion, sexual orientation, gender identity, national origin, status as a veteran, and basis of disability or any federal, state, or local protected class.

Additional Locations: * United States of America - North Carolina - Morrisville

If you require an accommodation to complete this application, please contact ability@lenovo.com

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cluster Solution Architect
Cluster Solution Architect

Lenovo • Morrisville

On-site
CAD 198,000 - 283,000
AI/HPC Product Line Manager
AI/HPC Product Line Manager

Lenovo • Morrisville

On-site
CAD 209,000 - 293,000
Sr. Technology Consultant- Hybrid Cloud
Sr. Technology Consultant- Hybrid Cloud

Lenovo • North Sydney

On-site
CAD 100,000 - 140,000
Principal Private Cloud Architect
Principal Private Cloud Architect

Lenovo • Morrisville

On-site
CAD 296,000 - 453,000
Junior Cloud Platform Engineer
Junior Cloud Platform Engineer

Lenovo • Morrisville

On-site
CAD 121,000 - 157,000
AI Application Developer
AI Application Developer

Lenovo • Morrisville

Hybrid
CAD 157,000 - 228,000
Cloud Native PaaS Platform Operations Staff Engineer
Cloud Native PaaS Platform Operations Staff Engineer

Lenovo • Morrisville

On-site
CAD 170,000 - 226,000
Storage Architect, Solutions & Ecosystem Strategy
Storage Architect, Solutions & Ecosystem Strategy

Lenovo • Morrisville

On-site
CAD 256,000 - 341,000
Storage Product Line Manager
Storage Product Line Manager

Lenovo • Morrisville

On-site
CAD 209,000 - 265,000
Data Center System Administrator
Data Center System Administrator

Lenovo • Montreal (administrative region)

On-site
CAD 67,000 - 103,000