The OCI AI Infrastructure Network Operations team operates and improves the high-performance RDMA/RoCE network fabrics powering OCI's largest AI, GPU, and HPC workloads.
As a Senior Manager, you will lead a team responsible for building, operating, and scaling these critical network fabrics and supporting systems. You will combine deep networking expertise in RDMA/RoCE, Clos fabrics, congestion control, telemetry, and performance troubleshooting with strong software engineering and people leadership.
You will drive automation, monitoring, resiliency, and operational readiness while partnering across Network Availability, Automation, Monitoring, GNOC, hardware engineering, and service teams. Your team will improve network performance and availability, resolve complex customer issues, and build fault-tolerant systems that support AI infrastructure at global cloud scale.
Internal Responsibilities
As a Senior Manager in the AI Infrastructure Network Operations organization, you will:
- Lead and develop a team of engineers responsible for RDMA/RoCE fabric operations, performance, automation, and troubleshooting across OCI's AI/HPC infrastructure.
- Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems supporting hyperscale workloads.
- Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting to improve network performance and availability.
- Guide the architecture and development of operational tools, automation platforms, monitoring systems, and infrastructure services, including Infrastructure as Code (IaC).
- Drive improvements in resiliency, observability, testing, and automation while simplifying and scaling operational workflows.
- Lead operational readiness, customer escalations, NOC events, and complex production incidents, coordinating resolution across networking, software, hardware, and operations teams.
- Define team roadmaps and data-driven KPIs focused on fabric health, engineering efficiency, operational backlog, customer impact, performance, and service availability.
- Partner with Network Availability, Network Automation, Network Monitoring, GNOC, deployment, hardware, and service teams to deliver reliable infrastructure at cloud scale.
- Ensure operational planning, security, compliance, change management, staffing, on-call coverage, and service-level expectations are met.
- Drive continuous improvement in engineering practices, processes, tooling, and operational efficiency.
- Attract, mentor, and develop engineers across networking, software development, automation, and distributed systems while building a high-performing engineering organization.
- Participate in the manager on-call rotation and provide technical and organizational leadership during high-severity incidents.
Preferred ExperienceStrong background in operating or building network for large-scale cloud.
Experience with RDMA/RoCE, GPU/HPC networking, Clos fabrics, congestion management, telemetry, and performance debugging.
External Responsibilities
As a Senior Manager in the AI Infrastructure Network Operations organization, you will:
- Lead and develop a team of engineers responsible for RDMA/RoCE fabric operations, performance, automation, and troubleshooting across OCI's AI/HPC infrastructure.
- Drive the design, operation, scalability, reliability, and performance of highly available network and distributed systems supporting hyperscale workloads.
- Apply deep expertise in RDMA, RoCE, Ethernet fabrics, congestion control, QoS, telemetry, and large-scale troubleshooting to improve network performance and availability.
- Guide the architecture and development of operational tools, automation platforms, monitoring systems, and infrastructure services, including Infrastructure as Code (IaC).
- Drive improvements in resiliency, observability, testing, and automation while simplifying and scaling operational workflows.
- Lead operational readiness, customer escalations, NOC events, and complex production incidents, coordinating resolution across networking, software, hardware, and operations teams.
- Define team roadmaps and data-driven KPIs focused on fabric health, engineering efficiency, operational backlog, customer impact, performance, and service availability.
- Partner with Network Availability, Network Automation, Network Monitoring, GNOC, deployment, hardware, and service teams to deliver reliable infrastructure at cloud scale.
- Ensure operational planning, security, compliance, change management, staffing, on-call coverage, and service-level expectations are met.
- Drive continuous improvement in engineering practices, processes, tooling, and operational efficiency.
- Attract, mentor, and develop engineers across networking, software development, automation, and distributed systems while building a high-performing engineering organization.
- Participate in the manager on-call rotation and provide technical and organizational leadership during high-severity incidents.
Preferred ExperienceStrong background in operating or building network for large-scale cloud.
Experience with RDMA/RoCE, GPU/HPC networking, Clos fabrics, congestion management, telemetry, and performance debugging.