Site Reliability Engineer, Apple Data Platform / Big Data Platform

Apple Inc.

Austin, Northern (TX, KY)

Hybrid

USD 120,000 - 170,000

Full time

13 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. in Austin, TX seeks a Site Reliability Engineer to own the operational health of Apple Data Platform services, spanning Spark, Flink, Airflow, Trino, and governance tooling. You’ll collaborate with internal customers, engineering teams, and on-call responders to deliver reliable multi-cloud infrastructure.

You’ll drive monitoring, alerting, and automation, with on-call duties and a focus on reducing toil while scaling reliability across the era of data and AI products at Apple.

Qualifications

  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python; working knowledge of Golang is a plus.
  • Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Excellent written and verbal communication skills; ability to explain technical issues to non-experts.
  • Solid grounding in SRE principles with on-call, production support, or customer-facing experience.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio.
  • Participate in a rotating on-call schedule across supported services.
  • Own the operational health of big data platform services as SME; driving reliability and customer guidance.
  • Serve as a primary contact for internal customers via Slack, communicating status during issues.
  • Screen, triage, and resolve customer-reported issues and tickets based on impact.
  • Collaborate with dev teams to onboard new services, design monitoring and dashboards.
  • Build automation and self-healing tooling to reduce manual toil.
  • Identify and resolve production issues to protect reliability and customer experience.
  • Collaborate with partner teams to align execution with goals.

Skills

Python
Kubernetes
AWS/GCP
Big Data (Spark, Flink, Airflow, Trino

Education

Bachelor's degree in CS or related field

Tools

Prometheus
Grafana
Splunk
REST Catalog (Glue Catalog)

Job description

Site Reliability Engineer, Apple Data Platform / Big Data Platform

Austin, Texas, United States Software and Services

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands‑on support to internal teams, and partnering with developers to make cutting‑edge services like Spark, Flink, Airflow, Trino, Notebooks, and LLM‑based agent platforms reliable at scale.

Description

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in the big data engines and catalog/governance layers that power analytics and data engineering across the company. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from ML/AI platform services to multi‑cloud infrastructure, and grow into the team's go‑to expert for big data platform services — including Spark, Flink, Airflow, Trino, Notebooks, REST Catalog services (such as Glue Catalog), and data governance. Just as importantly, you'll be a first point of contact for the internal customers who rely on these services daily — someone who can translate a confusing error or a vague support request into a clear diagnosis and a fast resolution. We're looking for a self‑motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love solving hard operational problems, take genuine satisfaction in helping frustrated customers get unblocked, and want a front‑row seat to how Apple's data engineering platform scales, this role offers real room to grow your scope and impact over time.

Responsibilities
  • Operate, monitor, and triage production and non-production environments across the ADP portfolio — data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of big data platform services as SME — driving reliability, support, and customer guidance for Spark, Flink, Airflow, Trino, Notebooks, REST Catalog, and governance tooling.
  • Serve as a primary point of contact for internal customers via Slack — clearly communicating status, root cause, and next steps during active issues.
  • Screen, triage, and resolve customer-reported service issues and support tickets, prioritizing based on customer impact and urgency.
  • Partner with dev teams across time zones to onboard new services — understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Build automation and self-healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, elevate, and resolve production issues to protect platform reliability and customer experience.
  • Champion customer success by helping internal teams understand platform capabilities and adopt tools effectively.
  • Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.
Minimum Qualifications
  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Excellent written and verbal communication skills, with the ability to explain technical issues clearly to non-expert customers.
  • Solid grounding in SRE principles, with prior on-call, production-support, or customer-facing support role experience.
Preferred Qualifications
  • Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.
  • Prior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.
  • Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
  • Working knowledge of CI/CD pipelines and deployment workflows.
  • Experience with S3 and cloud storage/networking fundamentals.
  • Familiarity with data pipeline orchestration and workflow scheduling patterns.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 140,000 - 180,000
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Site Reliability Engineer, AiDP Production Engineering
Site Reliability Engineer, AiDP Production Engineering

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000
Data Platform SRE, AI & Data Platforms (AiDP)
Data Platform SRE, AI & Data Platforms (AiDP)

Apple Inc. • Austin (TX)

On-site
USD 100,000 - 130,000