Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Apple Inc.

Austin, Northern (TX, KY)

Hybrid

USD 120,000 - 180,000

Full time

6 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. in Austin, Texas, seeks a Site Reliability Engineer to own and operate Apple Data Platform's multi-cloud infrastructure. You’ll manage data pipelines, ML/AI services, and Kubernetes clusters across AWS and on-prem environments, delivering reliability at scale.

You'll collaborate with internal engineers, implement monitoring and GitOps workflows using Flux and Crossplane, and drive improvements to reduce toil while expanding the platform’s capabilities for internal customers.

Qualifications

  • Bachelor's Degree in CS or eng or equivalent.
  • 1-4 years in SRE/DevOps/Infrastructure role.
  • Proficient in Python; Golang a plus.
  • Kubernetes admin: RBAC, scheduling, autoscalers, PDBs, cluster-wide troubleshooting.
  • Strong communication and composure during incidents.
  • Solid grounding in SRE principles with on-call/production support.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio.
  • Participate in rotating on-call schedule, including occasional weekday/weekend coverage.
  • Own operational health of multi-cloud infra as SME — AWS, EKS, cross-cloud networking.
  • Provide Slack-based support; screen, triage, and resolve issues.
  • Debug IAM permission errors, storage quotas, namespace separation, and cluster-wide disruptions.
  • Partner with dev teams to onboard new services; design monitoring, alerting, dashboards (Prometheus, Grafana, Splunk).
  • Maintain and evolve Infrastructure-as-Code (Crossplane, Terraform) and GitOps (Flux) workflows.
  • Build automation and self-healing tooling to reduce toil.
  • Identify, escalate, and resolve production issues to protect reliability.
  • Collaborate with SRE and dev partner teams to align execution with goals.

Skills

Python
Kubernetes
RBAC
GitOps
Crossplane
Terraform
Prometheus
Grafana
Splunk
Spark
Flink

Education

Bachelor's Degree in Computer Science

Tools

Crossplane
Terraform
Flux

Job description

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Austin, Texas, United States Software and Services

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands‑on support to internal teams, and partnering with developers to make cutting‑edge services like Spark, Flink, Airflow, Ray, Notebooks, and LLM‑based agent platforms reliable at scale across AWS, GCP, and on‑premise Kubernetes.

Description

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in the cloud infrastructure that underpins all of it. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from big data pipelines to ML/AI platform services, and grow into the team's go-to expert for multi‑cloud infrastructure — including AWS core services (IAM, EKS, RDS, S3, VPC networking, autoscaling, EBS), Kubernetes administration at scale, and the Infrastructure‑as‑Code and GitOps tooling that keeps it all reconciled and reliable. You'll be the person other engineers turn to when an IAM policy misfires, a cluster hits a scheduling wall, or a GitOps reconciliation drifts out of sync — and the driving force behind making those failure modes rarer over time.We're looking for a self‑motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love going deep on cloud and Kubernetes internals, enjoy being the trusted expert customers turn to, and want a front‑row seat to Apple Data Platform's multi‑cloud evolution, this role offers real room to grow your scope and impact over time.

Responsibilities
  • Operate, monitor, and triage production and non-production environments across the ADP portfolio — data processing, ML/AI, and multi‑cloud infrastructure.
  • Participate in a rotating on‑call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of multi‑cloud infrastructure as SME — driving reliability, support, and customer guidance for AWS services, EKS clusters, and cross‑cloud networking.
  • Provide Slack‑based support to internal customers; screen, triage, and resolve service related issues.
  • Debug production incidents involving IAM permission errors, storage quota limits, control‑plane/data‑plane namespace separation, and cluster‑wide disruptions.
  • Partner with dev teams across time zones to onboard new services — understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Maintain and evolve Infrastructure‑as‑Code (Crossplane, Terraform) and GitOps (Flux) workflows, troubleshooting state drift and reconciliation issues.
  • Build automation and self‑healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, escalat, and resolve production issues to protect platform reliability and customer experience.
  • Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.
Minimum Qualifications
  • Minimum Qualifications
  • * Bachelor's Degree in Computer Science, an engineering‑related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure‑focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Kubernetes administration experience — RBAC, node/pod scheduling, autoscalers, PriorityClasses/PDBs, and troubleshooting cluster‑wide disruptions.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on‑call or production‑support experience.
Preferred Qualifications
  • Experience with Infrastructure‑as‑Code (Crossplane and/or Terraform), including debugging state drift and composition/controller issues.
  • Experience with GitOps workflows (Flux or similar) — HelmRepository/reconciliation troubleshooting and Helm chart deployment.
  • Multi‑cloud exposure (GCP) — parity and migration scenarios are emerging areas of focus.
  • Experience with Splunk for log pipeline debugging (e.g., fluent‑bit).
  • Familiarity with Spark/Flink running on Kubernetes (executor scheduling, node affinity).
  • Comfort with GitHub PR review workflows in an infrastructure‑as‑code / GitOps context.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 140,000 - 180,000
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Site Reliability Engineer, AiDP Production Engineering
Site Reliability Engineer, AiDP Production Engineering

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Site Reliability Engineer (Edge Services), Infrastructure Services
Site Reliability Engineer (Edge Services), Infrastructure Services

Apple Inc. • Elk Grove (CA)

On-site
USD 132,000 - 245,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer - ASE / iCloud
Senior Site Reliability Engineer - ASE / iCloud

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 308,500
Medical and dental coverage
Retirement benefits
Stock programs (RSUs)
+2