Observability SRE Manager, Apple Services Engineering

Socket.dev

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Socket.dev seeks a senior engineering leader to set reliability strategy for our platform in Seattle, mentoring engineers, driving automation, and partnering with software, infra, finance and product teams to uphold reliability at scale.

You will shape SRE practices, observability, and AI-enabled tooling, requiring strategic vision and hands-on delivery at Apple-scale ambitions.

Qualifications

  • 5+ years engineering management experience leading SRE/infrastructure/observability teams
  • Experience hiring, mentoring, and building high-performing teams
  • Deep understanding of observability systems: metrics, logs, tracing, SLOs, error budgets
  • Proven ability to drive cross-functional initiatives to completion and communicate at executive level

Responsibilities

  • Define and execute the reliability strategy and roadmap for multi-tenant platforms
  • Mentor engineers, drive automation, and scale SRE practices
  • Lead cross-functional partnerships (software, infra, finance, product) to deliver scalable improvements
  • Build a culture of psychological safety and high performance within the team

Skills

Engineering management
SRE leadership
Observability
Cross-functional collaboration
Technical leadership
AI/ML tooling
Capacity planning
Root-cause analysis

Education

Bachelor's or Master's in CS/Engineering

Tools

Kubernetes
Prometheus
OpenTelemetry

Job description

DESCRIPTION

You'll set technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software, infrastructure, finance, and product teams to uphold the reliability of the platform whilst shipping improvements that matter at Apple scale. This is a senior leadership position that defines where our observability platform reliability, scalability and performance goes next, how our SRE practice evolves, including how AI reshapes it, and how we build the team and partnerships to get there. You will lead engineers solving reliability and scale problems few organizations encounter, integrating monitoring seamlessly across disparate infrastructures, hardware, software, application, and network layers, at a scale built to reach every user on the planet. You\'ll build a team culture that makes SRE sustainable, rewarding, and central to how Apple ships services. The successful candidate has a strong aptitude for both technical leadership and people management, with the ability to context-switch between strategic planning and tactical execution. You should be comfortable building and scaling teams, driving complex cross-functional initiatives, and thriving under pressure, while creating an inclusive, high-trust team culture where engineers do their best work. We believe AI will fundamentally reshape how SRE is practiced, from anomaly detection and root-cause analysis to capacity planning and toil elimination, and we\'re looking for a leader who shares that conviction and can drive that transformation across the organization.

MINIMUM QUALIFICATIONS

5+ years of engineering management experience leading SRE, infrastructure, or observability/monitoring teams Experience hiring and leading engineers, with a desire to build, grow, and mentor a team Deep understanding of observability systems and practices: metrics, logging, tracing, alerting, SLOs, error budgets, and fault analysis at scale Strong systems background, comfortable troubleshooting across the full stack (network, OS, container runtime, application) Experience operating large-scale, multi-tenant distributed systems in production, including Kubernetes environments Practical, solid knowledge of shell/bash scripting and at least one higher-level production language (Python preferred; Go, Java, or Scala also valued) Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations Track record of building high-performing teams through coaching, clear expectations, and psychological safety Demonstrated ability to drive cross-functional initiatives to completion and communicate at the executive level Bachelor\'s or Master\'s degree in Computer Science, Engineering, or related field, or equivalent experience

PREFERRED QUALIFICATIONS

Deep familiarity with the Prometheus ecosystem and cloud-native observability stacks (Thanos, Splunk, OpenTelemetry, or similar) Experience with third-party cloud platforms (AWS, GCP, or Azure) and infrastructure as code (Terraform, Ansible) Comfortable with open-source configuration management and orchestration tools (Helm, Puppet, Spinnaker) Demonstrable knowledge of TCP/IP, HTTP, web application security, and multi-tier web application architectures Experience running infrastructure as an internal managed service with defined SLAs Familiarity with microservices architecture and container orchestration with Kubernetes at scale Background in capacity planning, performance engineering, or infrastructure architecture Track record of driving cultural and process transformation within SRE organizations Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics) Developing and delivering multi-mode communications tailored to the unique needs of different audiences Anticipating and balancing the needs of multiple stakeholders Making sense of complex, high-quantity, and sometimes contradictory information to solve problems effectively Rebounding from setbacks and adversity when facing difficult situations Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) Manager, Apple Maps
Site Reliability Engineering (SRE) Manager, Apple Maps

Apple • Cupertino (CA)

On-site
USD 180,000 - 240,000
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple Services Engineering (ASE) Compute - Software Engineering Manager

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 240,000
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
SRE Engineering Program Manager: iCloud (ASE)
SRE Engineering Program Manager: iCloud (ASE)

Socket.dev • Seattle (WA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer (Edge Services), Infrastructure Services
Site Reliability Engineer (Edge Services), Infrastructure Services

Apple Inc. • Austin (TX)

On-site
USD 110,000 - 150,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
ASE Observability
ASE Observability

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 230,000
ASE Compute - Senior SRE Software Engineer
ASE Compute - Senior SRE Software Engineer

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000