Network Architect

GMI Cloud, Inc

United States

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

GMI Cloud is seeking an experienced Network Architect to join the Global Infrastructure team. You will lead the design, deployment, and day-to-day operations of high-performance network infrastructure supporting GPU-accelerated AI/ML workloads.

Preferred Location: US, Taipei, APAC. Responsibilities include designing InfiniBand/RoCE fabrics, planning global data-center networks, and driving automation with Ansible and REST APIs. Travel to data centers will be required.

Qualifications

  • Bachelor's degree in CS or related field.
  • 10+ years networking experience with 3+ years in InfiniBand/RoCE fabrics for large GPU clusters.
  • Hands-on InfiniBand HDR/XDR fabrics, subnet manager, routing, and P_Key knowledge.
  • Experience with RDMA, RoCE, and Ethernet lossless frameworks (PFC/ECN/DCB).
  • Strong troubleshooting: packet analysis, counters, link diagnostics, congestion root-cause.
  • Familiarity with NCCL, MPI, GPUDirect RDMA and GPU networking patterns.
  • Experience with NVMe-oF, NFS over RDMA, GPUDirect Storage.
  • Hands-on with Mellanox/NVIDIA Spectrum, Cisco hardware.
  • TCP/IP, VLANs, L3 routing, MTU/Jumbo frames, L2/L3 troubleshooting.
  • Monitoring tools (Prometheus, Grafana, SNMP, ELK) and network automation (Ansible, REST APIs).
  • Security knowledge: DDOS, IDS; cross-functional collaboration; bilingual in English/Chinese preferred.

Responsibilities

  • Architect, design, and implement high-performance network fabrics for GPU clusters.
  • Plan and deploy global data center networks, WAN, core, DC, firewalls and load balancers.
  • Build networks for Compute, Storage, Inband, Management, and OOB using InfiniBand and RoCEv2.
  • Lead hardware lifecycle: HBAs, NICs, RDMA NICs, switches and firmware.
  • Configure RDMA, QoS, ECN, PFC and Flow Control to optimize MPI/NCCL.
  • Operate day-to-day: provisioning, patches, upgrades, monitoring.
  • Troubleshoot production incidents, drive RCA, implement fixes.
  • Develop automation, runbooks, and capacity planning (Ansible, REST APIs).
  • Collaborate with compute, storage, platform, and SRE teams; validate end-to-end perf.
  • Identify providers and solutions; enforce network security and segmentation.
  • Maintain diagrams, cabling maps, baselines, and procedures; mentor juniors.
  • Occasional travel to GMI data centers.

Skills

InfiniBand/RoCE design
RDMA
Ethernet lossless frameworks
NCCL/MPI
GPU-communication patterns
Network troubleshooting
Automation scripting
Ansible
REST APIs
Prometheus
Grafana
SNMP
ELK
Mellanox/NVIDIA Spectrum
Cisco switches/routers
TCP/IP
BGP/OSPF
GPUDirect RDMA
Network security

Education

Bachelor’s degree in Computer Science or related field

Tools

Ansible
Prometheus
Grafana
SNMP
ELK
Mellanox/NVIDIA Spectrum
Cisco

Job description

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

About the Role

We are seeking an experienced Network Architect to join the GMI Global Infrastructure team. You will be responsible for leading the design, deployment, and day-to-day operations of high performance network infrastructure supporting GPU-accelerated AI/ML workloads. This role requires deep, hands-on expertise with InfiniBand and RoCE, high-throughput low-latency fabrics, network troubleshooting, performance tuning, and cross-functional collaboration with various stakeholders.

Preferred Location: US, Taipei, APAC.

Responsibilities
  • Architect and design high-performance, highly available network fabrics for GPU clusters (InfiniBand or RoCE), including topology, cabling, switch configuration, subnet/partitioning, and redundancy.
  • Plan, design, and implement network infrastructure for GMI global data center, including WAN, Core Network, Data Center Network, Firewalls, Load Balancers, DNS, VPN, etc.
  • Build high-performance network solutions to support AI/ML workloads, encompassing Compute, Storage, Inband, Management, and Out-of-Band (OOB) network fabric using Infiniband and Ethernet RDMA RoCEv2 technologies.
  • Lead implementation and lifecycle management of network hardware and firmware (HBAs, NICs, RDMA NICs, IB switches, Ethernet switches, etc).
  • Configure and fine-tune RDMA, congestion control, QoS, ECN, Priority Flow Control (PFC), and Flow Control to optimize MPI, NCCL, and other GPU-communication patterns.
  • Hands-on day-to-day operations: provisioning, configuration changes, patching, firmware upgrades, and logging/monitoring of network health.
  • Rapidly troubleshoot production incidents (fabric errors, performance degradations, link flaps, MTU/flow issues, packet drops), drive root cause analysis, and implement corrective actions.
  • Develop and maintain automation, scripts, and runbooks for deployment, configuration management, diagnostics, and capacity planning (Ansible, REST APIs).
  • Work with compute, storage, platform and SRE engineers to validate end-to-end performance, run benchmarks, and recommend architecture improvements; collaborate with cross-functional teams and stakeholders to understand networking requirements.
  • Identify suitable network providers, vendors, and solutions to meet organizational needs.
  • Define and enforce network security, segmentation, and access policies for GPU clusters and management networks.
  • Maintain documentation: network diagrams, cabling maps, configuration baselines, and operational procedures.
  • Mentor junior network engineers and participate in on-call rotations for fabric support.
  • Regional/international travel to GMI data center locations.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Qualifications
  • Bachelor’s degree in Computer Science or related field.
  • 10+ years of networking experience, with 3+ years specifically designing and operating InfiniBand/RoCE-based fabrics for large-scale GPU AI clusters (hundreds to thousands of GPUs).
  • Deep, hands-on experience with InfiniBand (HDR/XDR) fabrics: subnet manager (UFM, OpenSM), IB routing, partition keys (P_Key), etc.
  • Strong expertise with RDMA, RoCE, and Ethernet lossless frameworks: configuring PFC, ECN, DCB, and addressing head-of-line/blocking issues.
  • Proven troubleshooting experience: packet capture analysis, IB/Ethernet counters, link diagnostics, congestion root-cause, firmware and driver interactions.
  • Familiarity with GPU communication libraries and patterns: NCCL, MPI, GPUDirect RDMA, and how network settings affect scaling and latency.
  • Familiarity with storage protocols used in AI environments (NVMe-oF, NFS over RDMA, GPUDirect Storage).
  • Hands-on with network hardware from major vendors (Mellanox/NVIDIA Spectrum, Cisco, etc).
  • Solid understanding of TCP/IP, VLANs, L3 routing, BGP/OSPF basics, MTU/Jumbo frames, and L2/L3 troubleshooting tools.
  • Experience with monitoring and observability tools (Prometheus, Grafana, SNMP, ELK, vendor telemetry).
  • Experience with network automation and scripting (Ansible, REST API), and configuration management.
  • Familiar with various routers, switches, firewalls, load balancer, DNS, VPN configuration implementation.
  • Strong knowledge in network security, DDOS, IDS, etc.
  • Familiar with optical networking, including fibers, transceivers and optics troubleshooting.
  • Candidates holding network certifications (e.g. CCNA, CCNP, Nvidia) will be strongly preferred.
  • Candidates with proven experience in the AI/ML GPU networking environment will be highly considered.
  • Strong troubleshooting mindset: methodical, data-driven, and calm under production pressure.
  • Proactive about automation, reliability, and continuous improvement.
  • Bilingual English and Chinese will be strongly preferred.
  • Collaborative team player, able to work cross-functionally and to translate technical trade-offs for stakeholders with strong communication skills.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infra Network Operation Manager
Infra Network Operation Manager

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Network Operations Engineer
Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Infra Engineer - Network Operation
Infra Engineer - Network Operation

GMI Cloud, Inc • United States

On-site
USD 140,000 - 180,000
Senior Network Architect — AI GPU Infra (InfiniBand)
Senior Network Architect — AI GPU Infra (InfiniBand)

GMI Cloud, Inc • United States

On-site
USD 180,000 - 240,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior GPU Networking & Data Center Infra Engineer
Senior GPU Networking & Data Center Infra Engineer

GMI Cloud, Inc • United States

On-site
USD 140,000 - 180,000
Staff HPC Network Architect
Staff HPC Network Architect

Lambda • United States

Hybrid
USD 180,000 - 280,000
Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 120,000 - 190,000
Network Engineer
Network Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 180,000
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc
Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) - Radian Arc

Greenhouse Software, Inc. • United States

Remote
USD 180,000 - 240,000