Observability SRE Manager, Apple Services Engineering

Apple Inc.

Seattle (WA)

On-site

USD 226,000 - 338,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking an Observability SRE Manager for its Services Engineering organization in Seattle. You will lead reliability, monitoring, and self-service tooling across Apple’s observability stack, guiding AI-enhanced approaches to reduce toil and improve incidents handling.

The role requires managing a multi-tenant platform, defining roadmaps, and partnering with software, infrastructure, and product teams to ensure platform health at scale.

Qualifications

  • 5+ years of engineering management experience leading SRE, infrastructure, or observability/monitoring teams.
  • Deep understanding of observability systems: metrics, logging, tracing, alerting, SLOs, error budgets at scale.
  • Experience operating large-scale, multi-tenant distributed systems in production, including Kubernetes environments.

Responsibilities

  • Own reliability, availability, and performance of Apple’s observability platform and self-service capabilities.
  • Lead staging and production environments to maximize availability.
  • Define and drive strategic roadmap for observability infrastructure with SRE, engineering, and product stakeholders.
  • Establish SRE practices including SLOs, capacity planning, disaster recovery, and change management.
  • Drive incident response, post-incident reviews, and systemic improvements to reduce toil.
  • Hire, mentor, and grow a high-performing SRE team; build inclusive culture.

Skills

SRE management
Observability
Kubernetes
Python
AI/ML tooling

Education

Bachelor's or Master's in CS/Engineering

Tools

Prometheus
OpenTelemetry
Thanos
Splunk
AWS/GCP/Azure

Job description

Observability SRE Manager, Apple Services Engineering

Seattle, Washington, United States Software and Services

People at Apple don't just build products, they craft the kind of experiences that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here!Join Apple, and help us leave the world better than we found it.The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love.Apple's observability and monitoring platforms are the nervous system behind the reliability of Apple's cloud services, giving thousands of engineers the visibility they need to detect, diagnose, and resolve issues before they impact customers. Our cloud monitoring platform analyzes billions of metrics per minute and is the first place incident responders turn when something goes wrong, regardless of system scale or complexity.We're looking for a senior SRE leader to own and evolve this platform: the metrics, logging, tracing, and alerting infrastructure that underpins operational excellence across Apple Services Engineering.

Description

You'll set technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software, infrastructure, finance, and product teams to uphold the reliability of the platform whilst shipping improvements that matter at Apple scale.This is a senior leadership position that defines where our observability platform reliability, scalability and performance goes next, how our SRE practice evolves, including how AI reshapes it, and how we build the team and partnerships to get there. You will lead engineers solving reliability and scale problems few organizations encounter, integrating monitoring seamlessly across disparate infrastructures, hardware, software, application, and network layers, at a scale built to reach every user on the planet. You'll build a team culture that makes SRE sustainable, rewarding, and central to how Apple ships services.The successful candidate has a strong aptitude for both technical leadership and people management, with the ability to context-switch between strategic planning and tactical execution. You should be comfortable building and scaling teams, driving complex cross-functional initiatives, and thriving under pressure, while creating an inclusive, high-trust team culture where engineers do their best work.We believe AI will fundamentally reshape how SRE is practiced, from anomaly detection and root-cause analysis to capacity planning and toil elimination, and we're looking for a leader who shares that conviction and can drive that transformation across the organization.

Responsibilities
  • Technical & Operational Leadership
  • Own the reliability, availability, and performance of Apple's observability platform (metrics, logging, tracing, alerting) and the self-service capabilities built on top of it
  • Lead staging and production environments for the observability platform with the goal of maximizing availability
  • Ensure the platform accurately monitors the health of every application and piece of infrastructure across the Apple ecosystem, the "central nervous system" that other engineering teams and incident responders reach for first
  • Define and drive the strategic roadmap for observability infrastructure in partnership with SRE, engineering, and product stakeholders
  • Establish and refine SRE practices including SLOs, error budgets, capacity planning, scale testing, disaster recovery, and change management
  • Guide deep dives into systemic and latent reliability issues spanning the full stack (hardware, software, application, and network), partnering with software and systems engineers to drive fixes to resolution
  • Drive incident response, post-incident reviews, and systemic improvements that reduce operational toil
  • Champion automation to eliminate manual processes through tooling, self-service platforms, and APIs for internal customers
  • Manage on-call rotations and ensure sustainable, well-supported operational coverage
  • Drive standardization of monitoring and troubleshooting methodology across embedded SREs and services throughout the organization
  • Represent the SRE organization in design reviews and operational readiness exercises for new and existing services
  • Partner with software engineering and architecture teams to influence system design for reliability, scalability, and operability
  • Balance technical debt reduction with feature development to maintain platform health
  • AI & Modernization
  • Define and execute a clear AI strategy for the SRE organization, identifying high-impact opportunities where AI/ML tooling can reduce toil, accelerate root-cause analysis, and improve reliability outcomes
  • Drive adoption of AI-assisted tooling (copilots, intelligent runbooks, LLM-based diagnostics, anomaly detection) into day-to-day SRE workflows
  • Build a culture where engineers actively experiment with AI tools and modern approaches to solve operational problems
  • Lead, grow, and mentor a team of Site Reliability Engineers, conducting regular 1:1s, performance reviews, and career development discussions
  • Hire and build out the SRE team, with a desire to develop engineers to meet both their career goals and the organization's goals
  • Build and scale a high-performing team through coaching, clear expectations, and psychological safety
  • Foster an inclusive team culture and mentor diverse talent
  • Champion engineering best practices for code quality, system design, and operational excellence across the broader organization
  • Cross‑Functional & Executive Communication
  • Build strong partnerships across Apple, negotiating priorities and aligning on shared goals with an Apple‑first mindset
  • Represent the SRE perspective in cross‑functional planning, translating reliability requirements into architecture and investment decisions
  • Communicate effectively at the executive level, presenting strategy, trade‑offs, and progress to senior leadership
  • Drive standardization and best practices across the observability and infrastructure landscape
Minimum Qualifications
  • 5+ years of engineering management experience leading SRE, infrastructure, or observability/monitoring teams
  • Experience hiring and leading engineers, with a desire to build, grow, and mentor a team
  • Deep understanding of observability systems and practices: metrics, logging, tracing, alerting, SLOs, error budgets, and fault analysis at scale
  • Strong systems background, comfortable troubleshooting across the full stack (network, OS, container runtime, application)
  • Experience operating large‑scale, multi‑tenant distributed systems in production, including Kubernetes environments
  • Practical, solid knowledge of shell/bash scripting and at least one higher‑level production language (Python preferred; Go, Java, or Scala also valued)
  • Demonstrated experience applying AI/ML tooling or LLM‑based solutions to improve SRE or infrastructure operations
  • Track record of building high‑performing teams through coaching, clear expectations, and psychological safety
  • Demonstrated ability to drive cross‑functional initiatives to completion and communicate at the executive level
  • Bachelor's or Master's degree in Computer Science, Engineering, or related field, or equivalent experience
Preferred Qualifications
  • Deep familiarity with the Prometheus ecosystem and cloud‑native observability stacks (Thanos, Splunk, OpenTelemetry, or similar)
  • Experience with third‑party cloud platforms (AWS, GCP, or Azure) and infrastructure as code (Terraform, Ansible)
  • Comfortable with open‑source configuration management and orchestration tools (Helm, Puppet, Spinnaker)
  • Demonstrable knowledge of TCP/IP, HTTP, web application security, and multi‑tier web application architectures
  • Experience running infrastructure as an internal managed service with defined SLAs
  • Familiarity with microservices architecture and container orchestration with Kubernetes at scale
  • Background in capacity planning, performance engineering, or infrastructure architecture
  • Track record of driving cultural and process transformation within SRE organizations
  • Experience building or deploying AI‑powered operational tooling (AIOps, intelligent alerting, automated diagnostics)
  • Developing and delivering multi‑mode communications tailored to the unique needs of different audiences
  • Anticipating and balancing the needs of multiple stakeholders
  • Making sense of complex, high‑quantity, and sometimes contradictory information to solve problems effectively
  • Rebounding from setbacks and adversity when facing difficult situations
  • Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $225,600 and $338,400, and your base pay will depend on your skills, qualifications, experience, and location.
Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ASE Senior Site Reliability Engineer
ASE Senior Site Reliability Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
Site Reliability Engineer (Edge Services), Infrastructure Services
Site Reliability Engineer (Edge Services), Infrastructure Services

Apple Inc. • Denver (CO)

On-site
USD 132,000 - 245,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Cloud Monitoring SRE – job_id_000040 Job ID- 67
Cloud Monitoring SRE – job_id_000040 Job ID- 67

Apple • Pasadena (CA)

On-site
USD 115,000 - 204,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Site Reliability Engineering (SRE) Manager, Apple Maps
Site Reliability Engineering (SRE) Manager, Apple Maps

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 268,000 - 402,000
Service Reliability Engineer (SRE)
Service Reliability Engineer (SRE)

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 264,000
Medical and Dental coverage
Retirement benefits
Employee stock purchase plan
+2
ASE Observability
ASE Observability

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 175,000 - 263,000
Medical and dental coverage
Employee stock programs
Tuition reimbursement
+1
Site Reliability Engineer (Edge Services), Infrastructure Services
Site Reliability Engineer (Edge Services), Infrastructure Services

Apple Inc. • Elk Grove (CA)

On-site
USD 132,000 - 245,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

Apple Inc. • Sunnyvale (CA)

On-site
USD 216,200 - 324,800
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1