Get more replies from employers
Send a job-specific resume in minutes.
Apple Cloud Networking seeks an experienced Reliability Engineering Manager to lead a Production Engineering team responsible for the availability, latency, and resiliency of its global network services. You will mentor engineers, define standards, and partner with software, infrastructure and operations teams.
You will focus on SRE practices, release engineering, and data-driven decisions to prevent outages and improve operational maturity at scale.
Apple Cloud Networking team builds and operates large-scale, software-defined networking platforms that enable secure, resilient, and highly available multi-cloud connectivity with a global footprint. Our infrastructure powers critical Apple services, including iCloud, iTunes, Siri, and Maps. We are seeking an experienced and visionary Reliability Engineering Manager to lead and grow a team of engineers focused on ensuring the availability, performance, scalability, and resiliency of Apple’s global network services. In this role, you will work closely with software engineering, infrastructure, and operations teams across Apple to deliver reliable, fault-tolerant systems that operate at massive scale.
As a key leader within the Cloud Networking organization, you will define and drive the reliability and resiliency strategy for Apple’s network platform services. You will be responsible for building, scaling, and mentoring a high-performing Production Engineering team that champions SRE and SWE best practices, release engineering, and data-driven decision-making. You will establish strong cross-functional partnerships to ensure reliability and resiliency are embedded throughout the system lifecycle—from design and development to deployment and operations. Your leadership will help ensure Apple’s network services meet demanding availability, latency, resilience, and security requirements while continuously improving operational maturity. We are looking for a leader who is deeply passionate about operating mission-critical, globally distributed systems, preventing outages, learning from failures, and driving long-term reliability improvements.