AI-Driven GPU Network Operations Engineer

GMI Cloud

United States

On-site

USD 120,000 - 180,000

Full time

39 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GMI Cloud is seeking a Network Operation Engineer to join the Global Infrastructure team. The role focuses on GPU network infrastructure and advanced optical/data-center networking to support AI workloads across our worldwide data centers.

The candidate will design, build, monitor, and optimize high-performance networks, manage Infiniband and RDMA RoCEv2 technologies, and collaborate with cross-functional teams to ensure reliability and scalability of AI infrastructure.

Qualifications

  • Bachelor's degree in Computer Science or related field; 5+ years in senior network roles.
  • Strong problem-solving and communication skills; ability to troubleshoot complex networks.
  • Experience with GPU networking, data center fabrics, and high-performance networks.

Responsibilities

  • Plan, design, and implement network infrastructure for global data centers.
  • Build high-performance network solutions for AI/ML workloads (Compute, Storage, OOB).
  • Operate, monitor, and troubleshoot GPU/HPC networks (Infiniband, RoCe, Ethernet).
  • Manage optical transceivers and related components.
  • Collaborate with cross-functional teams to understand networking requirements.
  • Identify providers and vendors to meet organizational needs.
  • Ensure 99.9%+ network availability and implement security measures.
  • Document configurations, issues, and resolutions; conduct root-cause analysis.
  • Stay updated on optical communications and data center technologies.
  • Travel regionally/internationally to data center locations.

Skills

Problem solving
Communication
Team collaboration

Education

Bachelor's degree in CS or related field

Tools

Infiniband
RDMA RoCEv2
Routers and switches

Job description

GMI Cloud is seeking a Network Operation Engineer to join the Global Infrastructure team. The role focuses on GPU network infrastructure and advanced optical/data-center networking to support AI workloads across our worldwide data centers.

The candidate will design, build, monitor, and optimize high-performance networks, manage Infiniband and RDMA RoCEv2 technologies, and collaborate with cross-functional teams to ensure reliability and scalability of AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Network Operations Lead for AI Compute & Data Centers
GPU Network Operations Lead for AI Compute & Data Centers

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Network Operations Engineer
Network Operations Engineer

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Infra Network Operation Manager
Infra Network Operation Manager

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
AI Infra DevOps & Backend Engineer
AI Infra DevOps & Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
Site Reliability Lead
Site Reliability Lead

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
AI Cloud SRE Lead — Scale High-Performance GPU Infra
AI Cloud SRE Lead — Scale High-Performance GPU Infra

GMI Cloud • United States

On-site
USD 120,000 - 180,000