Senior Manager, RDMA Fabric Design & Engineering

Oracle

United States

On-site

USD 133,100 - 306,400

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical insurance
401(k) Matching
Stock Purchase Plan
Paid time off

Job summary

Oracle in the United States seeks a Senior Manager to lead design, architecture, and operations for large-scale RDMA backend fabrics powering AI, HPC, and cloud infrastructure. You will drive end-to-end network decisions and guide teams responsible for fabric validation, deployment, automation, reliability, and lifecycle management.

You will collaborate with silicon vendors, platform engineering, and cloud teams to deliver scalable, resilient networking solutions, establish SLOs and KPIs, and

Qualifications

  • 12+ years of experience in networking, distributed systems, cloud infrastructure, or network architecture.
  • 3+ years of people leadership experience with direct reports.
  • Deep understanding of Layer 2/Layer 3 networking, routing protocols, data center networking, and network architecture.
  • Experience designing or operating large-scale backend fabrics.

Responsibilities

  • Lead architecture and design of large-scale RDMA fabrics supporting AI training, inference, HPC, and cloud workloads.
  • Define network topology, routing strategy, congestion management, resiliency, and capacity models for multi-cluster deployments.
  • Drive technology evaluation and roadmap decisions across Ethernet, RoCE, InfiniBand, optical networking, and emerging fabric technologies.
  • Establish design standards, validation requirements, and deployment readiness criteria for production networks.
  • Partner with hardware, systems, and cloud engineering teams to ensure network architecture aligns with platform requirements.

Skills

RDMA Fabric Architecture
Cloud Networking
Data Center Networking
Network Scalability and Performance
Automation and Observability
People Leadership and Talent
Vendor and Stakeholder Management
Incident Management and Root Cause

Education

Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering

Tools

Python
Go
Ansible

Job description

Job Description

Leads the design, architecture, engineering, and operational strategy for large-scale RDMA backend fabrics supporting AI, HPC, and cloud infrastructure. This role is responsible for building and scaling high-performance, low-latency network fabrics, driving end‑to‑end network architecture decisions, and leading teams responsible for fabric design, validation, deployment, automation, reliability, and lifecycle management. The Senior Manager partners with silicon vendors, platform engineering, systems architecture, capacity planning, and cloud infrastructure teams to deliver scalable, resilient, and future‑ready networking solutions.

Key Responsibilities – RDMA Fabric Design, Architecture and Scale
  • Lead architecture and design of large-scale RDMA fabrics supporting AI training, inference, HPC, and cloud workloads.
  • Define network topology, routing strategy, congestion management, resiliency, and capacity models for multi‑cluster deployments.
  • Drive technology evaluation and roadmap decisions across Ethernet, RoCE, InfiniBand, optical networking, and emerging fabric technologies.
  • Establish design standards, validation requirements, and deployment readiness criteria for production networks.
  • Partner with hardware, systems, and cloud engineering teams to ensure network architecture aligns with platform requirements.
Network Engineering, Validation and Reliability
  • Oversee lab validation, scale testing, performance benchmarking, and failure scenario analysis.
  • Lead root cause analysis for complex network incidents involving congestion, packet loss, latency, or fabric instability.
  • Develop strategies to improve network availability, performance, and operational excellence.
  • Drive change management and deployment governance processes across multiple engineering teams.
  • Establish SLOs, KPIs, and reliability metrics for large‑scale backend networks.
Automation, Telemetry and Operations
  • Lead development of automation frameworks for deployment, configuration management, testing, and operations.
  • Drive adoption of telemetry, observability, anomaly detection, and network analytics platforms.
  • Champion infrastructure‑as‑code and software‑defined networking practices.
  • Improve operational efficiency through automation and self‑healing capabilities.
People Leadership and Cross‑Functional Collaboration
  • Build and lead a high‑performing team of network architects and engineers.
  • Coach and mentor engineers while developing future technical leaders.
  • Partner with product, cloud, systems, and infrastructure organizations to align business priorities with engineering execution.
  • Drive vendor engagement and influence roadmap decisions with strategic technology partners.
Minimum Qualifications
  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, or related discipline.
  • 12+ years of experience in networking, distributed systems, cloud infrastructure, or network architecture.
  • 3+ years of people leadership experience with direct reports.
  • Deep understanding of Layer 2/Layer 3 networking, routing protocols, data center networking, and network architecture.
  • Experience designing or operating large‑scale backend fabrics.
Preferred Qualifications
  • Experience with RDMA technologies including RoCEv2 and/or InfiniBand.
  • Experience building hyperscale cloud or large‑scale AI/HPC network infrastructure.
  • Knowledge of congestion control, ECMP, traffic engineering, QoS, and lossless Ethernet.
  • Experience with network automation using Python, Go, Ansible, or similar technologies.
  • Experience with telemetry platforms, network modeling, and capacity planning.
  • Experience working with public cloud providers or large cloud environments.
  • Advanced degree in Computer Science, Electrical Engineering, or related field.
Key Skills
  • RDMA Fabric Architecture
  • Cloud Networking
  • Data Center Networking
  • Network Scalability and Performance Engineering
  • Automation and Observability
  • People Leadership and Talent Development
  • Vendor and Stakeholder Management
  • Incident Management and Root Cause Analysis
Legal and Compliance

Disclaimer: Certain U.S. based or U.S. customer or client‑facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only.

US: Hiring Range in USD from: $133,100 - $306,400 per year. May be eligible for bonus, equity, and compensation deferral.

Compensation and Benefits
  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre‑tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non‑overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each calendar year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance

Career Level - M3

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Principal - Oracle Cloud Infrastructure AI Network Engineering
Senior Principal - Oracle Cloud Infrastructure AI Network Engineering

Oracle • Austin (TX)

On-site
USD 126,000 - 265,000
Health insurance
Disability insurance
Life insurance
+4
Senior Manager, AI Infrastructure Network Operations
Senior Manager, AI Infrastructure Network Operations

Oracle • Santa Clara (CA)

On-site
USD 133,000 - 307,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Flexible vacation
+1
Principal Network Development Engineer
Principal Network Development Engineer

Ll Oefentherapie • Seattle (WA)

On-site
USD 102,000 - 210,000
401(k) Savings and Investment Plan
Paid time off
Paid holidays
Network Developer 4
Network Developer 4

Oracle • Seattle (WA)

On-site
USD 102,000 - 210,000
Medical, dental & vision insurance
401(k) with company match
Paid time off & flexible vacation
Principal Network Engineer
Principal Network Engineer

Ll Oefentherapie • San Juan (PR)

On-site
USD 109,000 - 224,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan with company match
Paid parental leave
Principal Network Engineer
Principal Network Engineer

Oracle • Seattle (WA)

On-site
USD 109,000 - 224,000
Medical, dental, vision insurance
401(k) with company match
Employee stock purchase plan
+2
Principal Network Engineer
Principal Network Engineer

Oracle • San Jose (CA)

On-site
USD 109,000 - 224,000
Medical, dental, vision insurance
401(k) with company match
Paid time off, holidays, sick leave
+1
Lead Principal Network Engineer
Lead Principal Network Engineer

Oracle • United States

On-site
USD 146,000 - 306,000
Medical/dental/vision insurance
Disability coverage
401(k) match
+1
Lead Principal Network Engineer
Lead Principal Network Engineer

Oracle • Nashville (TN)

On-site
USD 146,000 - 306,000
Medical/Dental/Vision
Paid time off
401(k) plan
+1
Senior Network Developer 3 (IC3) – Network Region Build (NRB)
Senior Network Developer 3 (IC3) – Network Region Build (NRB)

Oracle • Nashville (TN)

On-site
USD 91,400 - 187,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Paid time off and holidays
+1