GPU Network Operations Lead for AI Compute & Data Centers

GMI Cloud

United States

On-site

USD 120,000 - 180,000

Full time

30 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

GMI Cloud is seeking a Network Operation Manager to join the Global Infrastructure team. The role focuses on GPU network infrastructure design, build, and optimization across data centers, with emphasis on Infiniband, RoCEv2, and optical components. Preferences include Denver or Taipei locations.

The candidate will monitor GPU/HPC networks, manage transceivers, and collaborate with cross-functional teams to meet demanding AI workloads and security requirements.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • Over 5 years of experience as senior network engineer or related position.
  • Extensive experience in network deployment and operational support.
  • Strong knowledge in network administration and architecture.
  • Strong knowledge in network security, DDOS, IDS, etc.
  • Familiar with routers, switches, firewalls, load balancers, DNS, VPN configuration.

Responsibilities

  • Plan, design, and implement network infrastructure for GMI global data center, including WAN, Core Network, Data Center Network, Firewalls, Load Balancers, DNS, VPN, etc.
  • Build high-performance network solutions to support AI/ML workloads using Infiniband and Ethernet RDMA RoCEv2; include Compute, Storage, Inband, Management, and Out-of-Band networking.
  • Operate, monitor and troubleshoot GPU/HPC network systems (Infiniband, RoCe, Ethernet) for optimal performance.
  • Manage and configure optical transceivers and related components.
  • Collaborate with cross-functional teams to understand business requirements related to networking.
  • Identify suitable network providers, vendors, and solutions to meet organizational needs.
  • Ensure high performance, scalability, and 99.9%+ availability of network services.
  • Implement network security measures and perform system upgrades.
  • Document network configurations, issues, and resolution procedures.
  • Conduct root cause analysis and resolve network outages efficiently.
  • Keep abreast of advancements in optical communications, GPU networking, and data center technologies.
  • Regional/international travel to GMI data center locations.

Skills

Senior engineer
Network design
Network operations
Network security
Troubleshooting
Communication

Education

Bachelor’s degree in Computer Science or related field

Tools

Routers
Switches
Firewalls
Load balancers
DNS
VPN

Job description

GMI Cloud is seeking a Network Operation Manager to join the Global Infrastructure team. The role focuses on GPU network infrastructure design, build, and optimization across data centers, with emphasis on Infiniband, RoCEv2, and optical components. Preferences include Denver or Taipei locations.

The candidate will monitor GPU/HPC networks, manage transceivers, and collaborate with cross-functional teams to meet demanding AI workloads and security requirements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-Driven GPU Network Operations Engineer
AI-Driven GPU Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Infra Network Operation Manager
Infra Network Operation Manager

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Network Operations Engineer
Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
GPU Network Architect for AI/HPC Clusters
GPU Network Architect for AI/HPC Clusters

CyberCoders • Santa Clara (CA)

Hybrid
USD 200,000 - 250,000
Health Benefits
401k
Relocation assistance
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Senior Manager GPU Cloud Networking & Infrastructure Equity
Senior Manager GPU Cloud Networking & Infrastructure Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 256,000 - 414,000
Technical Program Manager – AI Infrastructure / GPU Clusters
Technical Program Manager – AI Infrastructure / GPU Clusters

GMI Cloud • United States

On-site
USD 140,000 - 210,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
Site Reliability Lead
Site Reliability Lead

GMI Cloud • United States

On-site
USD 120,000 - 180,000