Site Reliability Engineer - Insights

Apple Inc.

Cupertino (CA)

On-site

USD 120,000 - 180,000

Full time

22 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Apple is seeking an extraordinary DevOps/SRE engineer to scale large data platforms and analytics tools within the Insight ecosystem. You will drive reliability, automation, and observability for a multi-exabyte data landscape supporting Apple’s products and manufacturing operations.

You’ll apply AI/ML techniques to operations, implement deep observability dashboards, and collaborate with global teams across time zones to improve system resilience and performance at scale.

Qualifications

  • 3+ years of experience in SRE, DevOps, or platform engineering.
  • Proficiency in Python or Java.
  • Experience with AWS or GCP and big-data or streaming tech like Kafka/Elasticsearch/Redis/Bigtable.
  • Experience applying AI/ML or LLM techniques to development or operations.

Responsibilities

  • Own reliability, performance, and scalability of services within the Insight ecosystem.
  • Drive automation to reduce toil and improve observability (SLOs/SLIs).
  • Lead incident response and post-incident reviews with architectural improvements.
  • Develop AI-driven alerting, anomaly detection, and automated incident triage.
  • Collaborate with cross-functional teams across Apple’s manufacturing services.

Skills

Python programming
Java programming
Analytical thinking

Education

BS/MS in Computer Science

Tools

Kafka
Elasticsearch
Redis
Bigtable
Druid
ClickHouse
Object storage
Docker
Kubernetes
Grafana
Prometheus
Kibana
ArgoCD
Jenkins
GitHub Actions

Job description

Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.Enterprise Technology Services (ETS) is part of IS&T and delivers global-scale platforms and services that keep Apple's operations secure and running. The team manages identity, device security, and anti-abuse platforms — covering everything from manufacturing and repairs to software updates and activations. ETS also oversees supply chain, manufacturing, and partner integration platforms, protecting data on more than 2.5 billion devices worldwide. And when Apple prepares for a global product launch, ETS owns the systems that ramp factory production — managing serial numbers, network credentials, and verified software.The Insight team runs one of Apple's most critical Big Data ecosystems — an exabyte-scale, highly-available infrastructure that underpins manufacturing operations for everyApple product, globally. Every iPhone, iPad, and Mac has touched our systems.We advance technology by relying on each other's strengths and skills to build somethingbigger than ourselves. For this reason, team culture is central to our values. We valuesocial skills and integrity as much as technical craft.

Description

We are looking for an extraordinary DevOps engineer with experience building large-scale data platforms, analytics tools, and solutions that can take our environment to thenext level. You thrive in a high-demand, fast-moving setting: you prioritize well, deliverahead of schedule, work independently, and raise the people around you.You will operate and improve services within a very large-scale, highly available Big Data ecosystem supporting Exabytes level of data with sustained, rapid growth. Your work directly enables the engineering and operations teams that build every Apple product.

Responsibilities
  • Own the reliability, performance, and scalability of services and infrastructure within the Insight ecosystem — with a build-to-manage mindset applied to design, deployment, and ongoing operations
  • Drive automation that reduces operational toil and improves ecosystem stability Instrument services for deep observability: dashboards, meaningful alerts, runbooks, and clearly defined SLOs and SLIs that give the team and stakeholders an accurate view of ecosystem health
  • Partner with incident management to drive effective incident response and conduct thorough post-incident reviews; translate findings into architectural and operational improvements that prevent recurrence
  • Build and advance AIOps capabilities across the ecosystem — developing AI- driven alerting, anomaly detection, LLM-assisted operational tooling, and automated incident triage that measurably improve observability and response
  • Collaborate with cross-functional engineering teams across Apple's global manufacturing services, communicating effectively across time zones and cultures
Minimum Qualifications
  • 3+ years of experience in SRE, DevOps, or platform engineering,with proficiency in Python or Java.
  • Experience with AWS or GCP, including big-data or streaming technologies such as Kafka, Elasticsearch, Redis, or Bigtable.
  • Experience applying AI/ML or LLM techniques to development or operations — such as anomaly detection, predictive alerting, or LLM/agent-driven tooling.
Preferred Qualifications
  • Experience with Distributed systems & big-data streaming: Microservices, APIs, and messaging/streaming (Kafka, Solace, Pub/Sub) at scale, plus production big-data/storage technologies (Druid, Elasticsearch, ClickHouse, or object storage).
  • Experience with Cloud-native delivery: Containerization and orchestration (Docker, Kubernetes)
  • and CI/CD or GitOps pipelines (ArgoCD, Jenkins, GitHub Actions, or similar).
  • Implementing and operating observability systems (Grafana, Prometheus, Kibana, or equivalent) instrumentation, dashboard design, and alert tuning.
  • BS/MS in Computer Science, Software Engineering, or equivalent experience; production relational databases (MySQL, PostgreSQL)multi-region operations and data-residency awareness; and a creative, flexible approach to problem-solving.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Site Reliability Engineer - AI & Data Platforms
Software Site Reliability Engineer - AI & Data Platforms

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 190,000
Sr. DevOps Engineer
Sr. DevOps Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 26,000 - 44,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

Apple Inc. • Cupertino (CA)

On-site
USD 94,000 - 141,000
Sr Site Reliability Engineer, Customer Systems , IS&T
Sr Site Reliability Engineer, Customer Systems , IS&T

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 210,000
Sr. Software Engineer (Data Solutions), IS&T Ai & Data Platforms
Sr. Software Engineer (Data Solutions), IS&T Ai & Data Platforms

Apple Inc. • Sunnyvale (CA), Northern (KY)

On-site
USD 185,000 - 278,000
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)

Apple Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 150,000 - 190,000
Observability SRE Manager, Apple Services Engineering
Observability SRE Manager, Apple Services Engineering

Apple Inc. • Seattle (WA)

On-site
USD 226,000 - 338,000
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Apple Inc. • Austin (TX), Northern (KY)

On-site
USD 120,000 - 180,000