Turn this role into an interview — a resume and cover letter built around what this employer wants.
Oracle Global Services Singapore Pte. Ltd. seeks a Principal Network Development Engineer to ensure reliability, scalability, and operational excellence of OCI networks. You will develop automation, analyze telemetry, and collaborate across teams to deliver highly available cloud services.
You will lead initiatives, mentor engineers, and contribute to runbooks and governance, while supporting incident response and service restoration in fast-paced, mission-critical environments.
As a Principal Network Development Engineer within Oracle Cloud Infrastructure (OCI) Network Reliability Engineering (NRE), you will help ensure the reliability, scalability, and operational excellence of one of the world's largest cloud networks. You will work on complex network infrastructure challenges, develop automation to improve operational efficiency, and collaborate across engineering and operations teams to deliver highly available cloud services.
Networking
- Expert-level knowledge of routing and networking protocols including BGP, OSPF, IS-IS, TCP/IP, IPv4, IPv6, DNS, DHCP, and MPLS.
- Strong hands-on experience with at least three of the following technologies:
- Proven ability to analyze network telemetry, identify root causes, and resolve complex service-impacting issues.
- Experience operating large-scale ISP, hyperscale cloud, or enterprise network environments.
- Familiarity with merchant silicon networking platforms such as Broadcom and Mellanox.
- Industry certifications (CCNP, CCIE, JNCIP, JNCIE, or equivalent) are advantageous.
GPU, RDMA & High-Performance Networking
- Experience supporting GPU-based infrastructure and AI/ML networking environments.
- Knowledge of RDMA technologies and lossless network architectures.
- Hands‑on experience with InfiniBand and High-Performance Computing (HPC) environments is highly desirable.
Automation & Software Development
- Develop and maintain automation solutions to improve operational efficiency and service reliability.
- Build scripts and tooling to automate routine operational tasks and troubleshooting workflows.
- Experience with one or more of the following:
Project & Technical Leadership
- Lead technical initiatives that improve operational processes, tooling, documentation, and service reliability.
- Drive the development and continuous improvement of runbooks, procedures, and operational standards.
- Contribute to strategic planning and execution of short, medium, and long-term engineering objectives.
- Partner with senior engineering and operational leaders to deliver critical infrastructure programs.
Preferred Experience & Attributes
- Extensive experience supporting large-scale network infrastructure in 24x7 production environments.
- Strong understanding of incident management, operational excellence, and service restoration practices.
- Excellent analytical, organizational, and problem‑solving skills.
- Strong written and verbal communication skills with the ability to influence technical and non‑technical stakeholders.
- Self‑motivated, proactive, and comfortable operating in fast-paced, mission‑critical environments.
- Willingness to participate in an on‑call or rotational support model as required.
Leadership Expectations
- Act as a technical leader within the organization, driving operational improvements and engineering best practices.
- Collaborate closely with shift leads, engineering teams, and management to ensure successful service delivery.
- Identify opportunities to improve team effectiveness through process optimization, automation, and innovation.
- Drive compliance reviews, runbook audits, and operational governance activities.
- Support recruitment, interviewing, mentoring, and development of junior engineers.
- Foster strong partnerships across OCI service teams to ensure consistent and scalable operational processes.