Infra Network Operation Manager

GMI Cloud

United States

On-site

USD 120,000 - 180,000

Full time

25 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

GMI Cloud is seeking a Network Operation Manager to join the Global Infrastructure team. The role focuses on GPU network infrastructure design, build, and optimization across data centers, with emphasis on Infiniband, RoCEv2, and optical components. Preferences include Denver or Taipei locations.

The candidate will monitor GPU/HPC networks, manage transceivers, and collaborate with cross-functional teams to meet demanding AI workloads and security requirements.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 5 years of experience as senior network engineer or related position.
  • Extensive experience in network deployment and operational support.
  • Strong knowledge in network administration and architecture.
  • Strong knowledge in network security, DDOS, IDS, etc.
  • Familiar with routers, switches, firewalls, load balancers, DNS, VPN configuration.

Responsibilities

  • Plan, design, and implement network infrastructure for GMI global data center, including WAN, Core Network, Data Center Network, Firewalls, Load Balancers, DNS, VPN, etc.
  • Build high-performance network solutions to support AI/ML workloads using Infiniband and Ethernet RDMA RoCEv2; include Compute, Storage, Inband, Management, and Out-of-Band networking.
  • Operate, monitor and troubleshoot GPU/HPC network systems (Infiniband, RoCe, Ethernet) for optimal performance.
  • Manage and configure optical transceivers and related components.
  • Collaborate with cross-functional teams to understand business requirements related to networking.
  • Identify suitable network providers, vendors, and solutions to meet organizational needs.
  • Ensure high performance, scalability, and 99.9%+ availability of network services.
  • Implement network security measures and perform system upgrades.
  • Document network configurations, issues, and resolution procedures.
  • Conduct root cause analysis and resolve network outages efficiently.
  • Keep abreast of advancements in optical communications, GPU networking, and data center technologies.
  • Regional/international travel to GMI data center locations.

Skills

Senior engineer
Network design
Network operations
Network security
Troubleshooting
Communication

Education

Bachelor’s degree in Computer Science or related field

Tools

Routers
Switches
Firewalls
Load balancers
DNS
VPN

Job description

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Network Operation Manager to join the GMI Global Infrastructure team. The ideal candidate will have hands-on experience with GPU network infrastructure and strong understanding of optical communication technologies. You will be responsible for designing, building, maintaining, monitoring, and optimizing advanced network systems to ensure high performance and reliability, supporting our global data centers and AI/ML workloads.

Preferred Location: Denver, Taipei.

Responsibilities
  • Plan, design, and implement network infrastructure for GMI global data center, including WAN, Core Network, Data Center Network, Firewalls, Load Balancers, DNS, VPN, etc.
  • Build high-performance network solutions to support AI/ML workloads, encompassing Compute, Storage, Inband, Management, and Out-of-Band (OOB) network fabric using Infiniband and Ethernet RDMA RoCEv2 technologies.
  • Operate, monitor and troubleshoot GPU/HPC network systems (Infiniband, RoCe, Ethernet) to ensure optimal performance on a daily basis.
  • Manage and configure optical transceivers and related components
  • Collaborate with cross-functional teams and stakeholders to understand business requirements related to networking
  • Identify suitable network providers, vendors, and solutions to meet organizational needs
  • Ensure high performance, scalability, and 99.9%+ availability of network services
  • Implement network security measures and perform system upgrades
  • Document network configurations, issues, and resolution procedures
  • Conduct root cause analysis and resolve network outages efficiently
  • Keep abreast of advancements in optical communications, GPU networking, and data center technologies
  • Regional/international travel to GMI data center locations.
Qualifications
  • Bachelor’s degree in Computer Science or related field.
  • Over 5 years of experience as senior network engineer or related position.
  • Extensive experience in network deployment and operational support.
  • Strong knowledge in network administration and architecture.
  • Strong knowledge in network security, DDOS, IDS, etc.
  • Familiar with various routers, switches, firewalls, load balancer, DNS, VPN configuration implementation.
  • Familiar with network monitoring tools and protocols.
  • Familiar with optical networking, including fibers, transceivers and optics.
  • Experience with high-performance data center networks, HPC or AI environments.
  • Excellent problem-solving, troubleshooting and communication skills.
  • Candidates holding network certifications (e.g. CCNA, CCNP) will be strongly preferred.
  • Candidates with proven experience in the AI/ML GPU networking environment will be highly considered.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Operations Engineer
Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
GPU Network Operations Lead for AI Compute & Data Centers
GPU Network Operations Lead for AI Compute & Data Centers

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Site Reliability Lead
Site Reliability Lead

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Technical Program Manager – AI Infrastructure / GPU Clusters
Technical Program Manager – AI Infrastructure / GPU Clusters

GMI Cloud • United States

On-site
USD 140,000 - 210,000
Sourcing Director of Strategic Procurement (Networking Infrastructure) - Global
Sourcing Director of Strategic Procurement (Networking Infrastructure) - Global

GMI Cloud • United States

On-site
USD 180,000 - 260,000
Network Engineer (AI GPU Cluster Operations)
Network Engineer (AI GPU Cluster Operations)

Aquila Hash, Inc. • Buffalo (NY)

On-site
USD 110,000 - 160,000