Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev

Austin (TX)

On-site

USD 140,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Data Platform SRE is seeking a self-motivated engineer to own multi-cloud infrastructure across AWS, GCP, and on-prem Kubernetes. You’ll operate and support the data platforms from data pipelines to ML/AI services, ensuring reliability at scale for internal engineers.

Join a team that values ownership, collaboration, and deep cloud & Kubernetes expertise, with opportunities to grow your scope as the platform evolves across Spark, Flink, and GitOps tooling.

Qualifications

  • Bachelor’s degree or equivalent experience in CS/engineering.
  • 1–4 years in SRE/DevOps/Infrastructure role.
  • Proficient in Python; Golang helpful.
  • Strong AWS experience: IAM, EKS, RDS, S3, VPC, autoscaling, EBS.
  • Kubernetes administration: RBAC, scheduling, autoscalers, PDBs, troubleshooting.
  • Strong incident response & on-call experience.
  • Solid grounding in SRE principles.

Responsibilities

  • Operate and support the Apple Data Platform SRE portfolio across multi-cloud infrastructure.
  • Manage cloud infrastructure across AWS, Kubernetes, and on-prem environments.
  • Provide on-call production support and incident response to internal teams.
  • Collaborate with engineers to improve reliability and resiliency of services.

Skills

Python
Golang
AWS
Kubernetes
Incident response
GitOps

Education

Bachelor's in CS/Engineering

Tools

Crossplane
Terraform
Flux
Helm
Splunk
Spark
Flink
Kubernetes

Job description

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books - at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries. Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success - running incident response, providing hands‑on support to internal teams, and partnering with developers to make cutting‑edge services like Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms reliable at scale across AWS, GCP, and on‑premise Kubernetes.

DESCRIPTION

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple - while specialising in the cloud infrastructure that underpins all of it. As an SRE on Apple Data Platform, you’ll operate and support the team's full portfolio, from big data pipelines to ML/AI platform services, and grow into the team's go‑to expert for multi‑cloud infrastructure - including AWS core services (IAM, EKS, RDS, S3, VPC networking, autoscaling, EBS), Kubernetes administration at scale, and the Infrastructure‑as‑Code and GitOps tooling that keeps it all reconciled and reliable. You’ll be the person other engineers turn to when an IAM policy misfires, a cluster hits a scheduling wall, or a GitOps reconciliation drifts out of sync - and the driving force behind making those failure modes rarer over time. We're looking for a self‑motivated engineer who thrives on ownership - someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love going deep on cloud and Kubernetes internals, enjoy being the trusted expert customers turn to, and want a front‑row seat to Apple Data Platform's multi‑cloud evolution, this role offers real room to grow your scope and impact over time.

MINIMUM QUALIFICATIONS
  • Minimum Qualifications * Bachelor's Degree in Computer Science, an engineering‑related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure‑focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Strong hands‑on AWS experience: IAM (roles, policies, permission boundaries, KMS), EKS, RDS, S3, VPC networking/endpoints, autoscaling groups, EBS.
  • Kubernetes administration experience - RBAC, node/pod scheduling, autoscalers, PriorityClasses/PDBs, and troubleshooting cluster‑wide disruptions.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on‑call or production‑support experience.
PREFERRED QUALIFICATIONS
  • Experience with Infrastructure‑as‑Code (Crossplane and/or Terraform), including debugging state drift and composition/controller issues.
  • Experience with GitOps workflows (Flux or similar) - HelmRepository/reconciliation troubleshooting and Helm chart deployment.
  • Multi‑cloud exposure (GCP) - parity and migration scenarios are emerging areas of focus.
  • Experience with Splunk for log pipeline debugging (e.g., fluent‑bit).
  • Familiarity with Spark/Flink running on Kubernetes (executor scheduling, node affinity).
  • Comfort with GitHub PR review workflows in an infrastructure‑as‑code / GitOps context.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning - for yourself, your team, and the org.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple Services Engineering (ASE) Compute - Software Engineering Manager

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 240,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000
Senior DevOps Engineer, Infrastructure Services Apple
Senior DevOps Engineer, Infrastructure Services Apple

Quest Technology Management • Elk Grove (CA)

Hybrid
USD 140,000 - 190,000
ASE Compute - Senior SRE Software Engineer
ASE Compute - Senior SRE Software Engineer

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Software Engineer, Cloud Services ASE
Software Engineer, Cloud Services ASE

Socket.dev • Austin (TX)

On-site
USD 180,000 - 240,000
SRE Engineering Program Manager: iCloud (ASE)
SRE Engineering Program Manager: iCloud (ASE)

Socket.dev • Seattle (WA)

On-site
USD 140,000 - 210,000
Senior DevOps Engineer, Infrastructure Services
Senior DevOps Engineer, Infrastructure Services

Socket.dev • California (MO)

On-site
USD 140,000 - 210,000