Network Operations Engineer

GMI Cloud

United States

On-site

USD 120,000 - 180,000

Full time

29 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

GMI Cloud is seeking a Network Operation Engineer to join the Global Infrastructure team. The role focuses on GPU network infrastructure and advanced optical/data-center networking to support AI workloads across our worldwide data centers.

The candidate will design, build, monitor, and optimize high-performance networks, manage Infiniband and RDMA RoCEv2 technologies, and collaborate with cross-functional teams to ensure reliability and scalability of AI infrastructure.

Qualifications

  • Bachelor's degree in Computer Science or related field; 5+ years in senior network roles.
  • Strong problem-solving and communication skills; ability to troubleshoot complex networks.
  • Experience with GPU networking, data center fabrics, and high-performance networks.

Responsibilities

  • Plan, design, and implement network infrastructure for global data centers.
  • Build high-performance network solutions for AI/ML workloads (Compute, Storage, OOB).
  • Operate, monitor, and troubleshoot GPU/HPC networks (Infiniband, RoCe, Ethernet).
  • Manage optical transceivers and related components.
  • Collaborate with cross-functional teams to understand networking requirements.
  • Identify providers and vendors to meet organizational needs.
  • Ensure 99.9%+ network availability and implement security measures.
  • Document configurations, issues, and resolutions; conduct root-cause analysis.
  • Stay updated on optical communications and data center technologies.
  • Travel regionally/internationally to data center locations.

Skills

Problem solving
Communication
Team collaboration

Education

Bachelor's degree in CS or related field

Tools

Infiniband
RDMA RoCEv2
Routers and switches

Job description

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents. Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter. From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud. One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Network Operation Engineer to join the GMI Global Infrastructure team. The ideal candidate will have hands‑on experience with GPU network infrastructure and strong understanding of optical communication technologies. You will be responsible for designing, building, maintaining, monitoring, and optimizing advanced network systems to ensure high performance and reliability, supporting our global data centers and AI/ML workloads.

Responsibilities
  • Plan, design, and implement network infrastructure for GMI global data center, including WAN, DC Core Network, Data Center Network, Firewalls, Load Balancers, DNS, VPN, etc.
  • Build high-performance network solutions to support AI/ML workloads, encompassing Compute, Storage, Inband, Management, and Out-of-Band (OOB) network fabric using Infiniband and Ethernet RDMA RoCEv2 technologies.
  • Operate, monitor and troubleshoot GPU/HPC network systems (Infiniband, RoCe, Ethernet) to ensure optimal performance on a daily basis.
  • Manage and configure optical transceivers and related components
  • Collaborate with cross-functional teams and stakeholders to understand business requirements related to networking
  • Identify suitable network providers, vendors, and solutions to meet organizational needs
  • Ensure high performance, scalability, and 99.9%+ availability of network services
  • Implement network security measures and perform system upgrades
  • Document network configurations, issues, and resolution procedures
  • Conduct root cause analysis and resolve network outages efficiently
  • Keep abreast of advancements in optical communications, GPU networking, and data center technologies
  • Regional/international travel to GMI data center locations.
Qualifications
  • Bachelor's degree in Computer Science or related field.
  • Over 5 years of experience as senior network engineer or related position.
  • Extensive experience in network deployment and operational support.
  • Strong knowledge in network administration and architecture.
  • Strong knowledge in network security, DDOS, IDS, etc.
  • Familiar with various routers, switches, firewalls, load balancer, DNS, VPN configuration implementation.
  • Familiar with network monitoring tools and protocols.
  • Familiar with optical networking, including fibers, transceivers and optics.
  • Experience with high-performance data center networks, HPC or AI environments.
  • Excellent problem-solving, troubleshooting and communication skills.
  • Candidates holding network certifications (e.g. CCNA, CCNP) will be strongly preferred.
  • Candidates with proven experience in the AI/ML GPU networking environment will be highly considered.

Meeting every qualification is not required - if you're excited about this role, we'd love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra Network Operation Manager
Infra Network Operation Manager

GMI Cloud • United States

On-site
USD 120,000 - 180,000
AI-Driven GPU Network Operations Engineer
AI-Driven GPU Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
Site Reliability Lead
Site Reliability Lead

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
GPU Network Operations Lead for AI Compute & Data Centers
GPU Network Operations Lead for AI Compute & Data Centers

GMI Cloud • United States

On-site
USD 120,000 - 180,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Network Engineer (AI GPU Cluster Operations)
Network Engineer (AI GPU Cluster Operations)

Aquila Hash, Inc. • Buffalo (NY)

On-site
USD 110,000 - 160,000
Network Engineer
Network Engineer

asobbi • United States

On-site
USD 160,000 - 190,000