Senior Manager, Core Infrastructure Engineering

Oracle Corporation

Nashville (TN)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oracle Corporation in Nashville seeks a Senior Manager to lead the AI Infrastructure Network Operations team, building, operating, and scaling RDMA/RoCE fabrics for large AI/HPC workloads.

You will drive automation, observability, and reliability across fabric operations, work with cross-functional teams, and mentor engineers to scale cloud-scale infrastructure.

Strong leadership, deep networking expertise in RDMA/RoCE, Clos fabrics, telemetry, and IaC experience are essential.

Qualifications

  • Lead a team of engineers responsible for RDMA/RoCE fabric operations, performance, automation, and troubleshooting across OCI's AI/HPC infrastructure.
  • Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems supporting hyperscale workloads.
  • Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting to improve network performance and availability.
  • Guide the architecture and development of operational tools, automation platforms, monitoring systems, and infrastructure services, including Infrastructure as Code (IaC).
  • Drive improvements in resiliency, observability, testing, and automation while simplifying and scaling operational workflows.
  • Lead operational readiness, customer escalations, NOC events, and complex production incidents, coordinating resolution across networking, software, hardware, and operations teams.
  • Define team roadmaps and data-driven KPIs focused on fabric health, engineering efficiency, operational backlog, customer impact, performance, and service availability.

Responsibilities

  • Lead and develop a team of engineers for RDMA/RoCE fabric operations, performance, automation, and troubleshooting.
  • Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems.
  • Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting.
  • Guide the architecture and development of operational tools, automation platforms, monitoring systems, and IaC-based infrastructure services.
  • Drive improvements in resiliency, observability, testing, and automation while scaling operational workflows.
  • Lead readiness, incidents, and coordination across networking, software, hardware, and operations teams.
  • Define roadmaps and KPIs focused on fabric health, engineering efficiency, and service availability.

Skills

RDMA RoCE
Clos fabrics
Network automation
Telemetry
Performance debugging

Tools

IaC

Job description

The OCI AI Infrastructure Network Operations team operates and improves the high-performance RDMA/RoCE network fabrics powering OCI's largest AI, GPU, and HPC workloads.

As a Senior Manager, you will lead a team responsible for building, operating, and scaling these critical network fabrics and supporting systems. You will combine deep networking expertise in RDMA/RoCE, Clos fabrics, congestion control, telemetry, and performance troubleshooting with strong software engineering and people leadership.

You will drive automation, monitoring, resiliency, and operational readiness while partnering across Network Availability, Automation, Monitoring, GNOC, hardware engineering, and service teams. Your team will improve network performance and availability, resolve complex customer issues, and build fault-tolerant systems that support AI infrastructure at global cloud scale.

Internal Responsibilities

As a Senior Manager in the AI Infrastructure Network Operations organization, you will:

  • Lead and develop a team of engineers responsible for RDMA/RoCE fabric operations, performance, automation, and troubleshooting across OCI's AI/HPC infrastructure.
  • Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems supporting hyperscale workloads.
  • Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting to improve network performance and availability.
  • Guide the architecture and development of operational tools, automation platforms, monitoring systems, and infrastructure services, including Infrastructure as Code (IaC).
  • Drive improvements in resiliency, observability, testing, and automation while simplifying and scaling operational workflows.
  • Lead operational readiness, customer escalations, NOC events, and complex production incidents, coordinating resolution across networking, software, hardware, and operations teams.
  • Define team roadmaps and data-driven KPIs focused on fabric health, engineering efficiency, operational backlog, customer impact, performance, and service availability.
  • Partner with Network Availability, Network Automation, Network Monitoring, GNOC, deployment, hardware, and service teams to deliver reliable infrastructure at cloud scale.
  • Ensure operational planning, security, compliance, change management, staffing, on-call coverage, and service-level expectations are met.
  • Drive continuous improvement in engineering practices, processes, tooling, and operational efficiency.
  • Attract, mentor, and develop engineers across networking, software development, automation, and distributed systems while building a high-performing engineering organization.
  • Participate in the manager on-call rotation and provide technical and organizational leadership during high-severity incidents.
Preferred Experience
  • Strong background in operating or building network for large-scale cloud.

  • Experience with RDMA/RoCE, GPU/HPC networking, Clos fabrics, congestion management, telemetry, and performance debugging.

External Responsibilities

As a Senior Manager in the AI Infrastructure Network Operations organization, you will:

  • Lead and develop a team of engineers responsible for RDMA/RoCE fabric operations, performance, automation, and troubleshooting across OCI's AI/HPC infrastructure.
  • Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems supporting hyperscale workloads.
  • Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting to improve network performance and availability.
  • Guide the architecture and development of operational tools, automation platforms, monitoring systems, and infrastructure services, including Infrastructure as Code (IaC).
  • Drive improvements in resiliency, observability, testing, and automation while simplifying and scaling operational workflows.
  • Lead operational readiness, customer escalations, NOC events, and complex production incidents, coordinating resolution across networking, software, hardware, and operations teams.
  • Define team roadmaps and data-driven KPIs focused on fabric health, engineering efficiency, operational backlog, customer impact, performance, and service availability.
  • Partner with Network Availability, Network Automation, Network Monitoring, GNOC, deployment, hardware, and service teams to deliver reliable infrastructure at cloud scale.
  • Ensure operational planning, security, compliance, change management, staffing, on-call coverage, and service-level expectations are met.
  • Drive continuous improvement in engineering practices, processes, tooling, and operational efficiency.
  • Attract, mentor, and develop engineers across networking, software development, automation, and distributed systems while building a high-performing engineering organization.
  • Participate in the manager on-call rotation and provide technical and organizational leadership during high-severity incidents.
Preferred Experience
  • Strong background in operating or building network for large-scale cloud.

  • Experience with RDMA/RoCE, GPU/HPC networking, Clos fabrics, congestion management, telemetry, and performance debugging.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Principal Engineer - AI Networking
Senior Principal Engineer - AI Networking

Oracle Corporation • Seattle (WA)

On-site
USD 180,000 - 260,000
Senior Network Developer - Topology Scaling
Senior Network Developer - Topology Scaling

Oracle • Albuquerque (NM)

On-site
USD 140,000 - 190,000
Senior Network Developer - Topology Scaling
Senior Network Developer - Topology Scaling

Oracle • Nashville (TN)

On-site
USD 140,000 - 200,000
Lead Principal Software Engineer, Core Infrastructure
Lead Principal Software Engineer, Core Infrastructure

Oracle Corporation • Seattle (WA)

On-site
USD 180,000 - 240,000
Network Developer 4
Network Developer 4

Oracle • Seattle (WA)

On-site
USD 102,300 - 209,500
Medical, dental & vision insurance
401(k) with company match
Paid time off & flexible vacation
Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle Corporation • Seattle (WA)

On-site
USD 150,000 - 210,000
Senior Manager, AI Infrastructure Networking & Automation
Senior Manager, AI Infrastructure Networking & Automation

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Ll Oefentherapie • Nashville (TN)

On-site
USD 150,000 - 180,000
Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Oracle Corporation • Reston (VA)

On-site
USD 130,000 - 190,000
Program Manager 5
Program Manager 5

Oracle • Nashville (TN)

On-site
USD 140,000 - 190,000