RDMA Ops Engineer - Computing Infrastructure Networking

Alibaba Cloud

Sunnyvale (CA)

On-site

USD 104,400 - 171,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading cloud service provider is seeking an RDMA Ops Engineer to optimize high-performance networking infrastructure for computing clusters. Responsibilities include deploying RDMA-based network architectures and optimizing performance. The ideal candidate has strong scripting skills, experience with RDMA technologies, and a solid understanding of Linux internals. This position offers a competitive salary range between $104,400 and $171,000/year based on market location and experience.

Qualifications

  • Strong scripting skills for operational automation.
  • Expert-level RDMA operational experience.
  • Understanding of Linux internals and proficient in tuning.

Responsibilities

  • Deploy and maintain RDMA-based network architectures.
  • Optimize network performance for distributed workloads.
  • Collaborate with algorithm engineers to troubleshoot network issues.

Skills

Scripting skills (Python/Go/Bash)
RDMA operational experience (RoCEv2/InfiniBand)
Linux network stack tuning
Network protocols (TCP/IP, RoCEv2)
Communication skills

Tools

Kubernetes networking (CNI, Multus, SR-IOV)
Automation tools

Job description

Overview

We're seeking a skilled RDMA Ops Engineer to optimize and maintain high-performance networking infrastructure for our computing clusters. This role focuses on building and operating ultra-low latency, high-throughput networks using RDMA technologies to power next-generation computing workloads.

Responsibilities
  • Deploy, operate and maintain RDMA-based network architectures (RoCE/InfiniBand) for cluster with thousands of nodes
  • Optimize network performance for distributed collective communication workloads (NCCL, MPI, etc.)
  • Solve complex network issues in distributed collective communication (e.g., NCCL/MPI communication bottlenecks)
  • Use automation tools for network provisioning, monitoring, diagnostics, and network performance profiling (latency/throughput analysis)
  • Implement CI/CD pipelines for network infrastructure-as-code
  • Manage end-to-end network lifecycle: deployment, configuration, monitoring, upgrades
  • Collaborate with computing algorithm engineers to troubleshoot network-related bottlenecks in training/inference pipelines
  • Bridge Computing framework requirements with underlying network infrastructure capabilities
  • Ensure compliance with security and scalability requirements
Qualifications
  • Strong scripting skills (Python/Go/Bash) for operational automation
  • Expert-level RDMA operational experience (RoCEv2/InfiniBand)
  • Understanding of Linux internals (kernel bypass, syscall optimization, etc), and proficient in Linux network stack tuning (irqbalance, NUMA, hugepages)
  • Hands-on experience with RDMA/DPDK performance tuning
  • Strong knowledge of network protocols (TCP/IP, RoCEv2) and NIC architecture principles
  • Ability to abstract complex technical concepts into architectural diagrams
  • Proven track record of translating R&D innovations into production solutions
  • Strong communication skills for cross-functional collaboration with Computing researchers and SRE teams
  • Experience managing production computing networks
  • Familiar with Kubernetes networking (CNI, Multus, SR-IOV) and GPU-aware scheduling
  • Background in computing system optimization (NVIDIA collective libraries, MPI tuning)
  • Deep understanding of computing workload patterns and their network implications
Compensation and Employment

The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. If hired, employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RDMA Ops Engineer for Ultra-Low Latency Clusters
RDMA Ops Engineer for Ultra-Low Latency Clusters

Alibaba Cloud • Sunnyvale (CA)

On-site
USD 104,000 - 171,000
Principal Engineer - AI Networking
Principal Engineer - AI Networking

Oracle • United States

On-site
USD 99,000 - 235,000
Potential bonus
Equity options
Compensation deferral
Senior Manager, RDMA Fabric Design & Engineering
Senior Manager, RDMA Fabric Design & Engineering

Oracle • United States

On-site
USD 133,000 - 307,000
Medical insurance
401(k) Matching
Stock Purchase Plan
+1
Principal Engineer - AI Networking
Principal Engineer - AI Networking

Oracle • Seattle (WA)

On-site
USD 114,000 - 235,000
Medical, dental, and vision insurance
401(k) with company match
Flexible vacation policy
+1
Transport (RDMA) Silicon Product Architect
Transport (RDMA) Silicon Product Architect

Intel • Santa Clara (CA)

Hybrid
USD 221,000 - 312,000
Stock bonuses
Health benefits
Retirement plans
+1
Senior Network Engineer: RDMA/RoCE, Automation & Global Ops
Senior Network Engineer: RDMA/RoCE, Automation & Global Ops

Oracle • San Jose (CA)

On-site
USD 109,000 - 224,000
Medical, dental, vision insurance
401(k) with company match
Paid time off, holidays, sick leave
+1
Senior Network Engineer: RDMA/RoCE, Automation & Global Ops
Senior Network Engineer: RDMA/RoCE, Automation & Global Ops

Oracle • Seattle (WA)

On-site
USD 109,000 - 224,000
Medical, dental, vision insurance
401(k) with company match
Employee stock purchase plan
+2
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • Houston (TX)

On-site
USD 150,000 - 240,000
Equity
Medical, dental, vision
Flexible paid time off
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • San Francisco (CA)

On-site
USD 150,000 - 240,000
Base + equity
Medical, dental, vision
Flexible PTO
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • Seattle (WA)

On-site
USD 150,000 - 240,000
Equity
Medical/dental/vision
Flexible PTO
+2