Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple

Austin (TX)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking a self‑motivated SRE for the Data Platform in Austin. You’ll operate and support a large multi‑cloud data platform powering internal AI and data services, focusing on reliability, incident response, and platform observability.

You’ll own the reliability roadmap, collaborate with cross‑functional teams, and help deploy scalable ML/AI pipelines and services (Ray training/serving, embeddings, vector stores) while staying aligned with Apple’s goals.

Qualifications

  • Bachelor's Degree in Computer Science or related field.
  • 1–4 years in a Site Reliability Engineering, DevOps, or Infrastructure‑focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Exposure to operating or supporting ML pipelines, model‑serving infrastructure, or LLM‑based systems in production.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on‑call or production‑support experience.

Responsibilities

  • Operate and support Apple's multi-cloud data platform and pipelines in production.
  • Lead incident response and drive reliability improvements across services.
  • Collaborate with internal teams to enable scalable ML/AI platform services (Ray training/serving, embeddings, vector stores).
  • Own reliability roadmap and advocate best practices across the organization.

Skills

Python
Golang
Kubernetes
AWS
GCP
ML pipelines
LLM infra
Production SRE

Education

Bachelor's degree in Computer Science or related field

Tools

Prometheus
Grafana
CI/CD pipelines
Ray
Notebooks

Job description

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands‑on support to internal teams, and partnering with developers to make cutting‑edge services like Spark, Flink, Airflow, Ray, Notebooks and LLM‑based agent platforms reliable at scale.

Description

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in an area that's shaping the future of how Apple builds and operates AI. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from big data pipelines to multi‑cloud infrastructure, and grow into the team's go‑to expert for ML/AI platform services — including Ray training and serving, LangGraph agent deployments, RAG architectures, embeddings platforms, and vector store platforms. You won't be building the models yourself, but you'll be the infrastructure backbone behind the teams who do — keeping their services, pipelines, and platforms running flawlessly in production so they can focus on innovation.

We're looking for a self‑motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love solving hard operational problems, enjoy being the trusted expert customers turn to, and want a front‑row seat to Apple's ML/AI infrastructure evolution, this role offers real room to grow your scope and impact over time.

Minimum Qualifications
  • Bachelor's Degree in Computer Science, an engineering‑related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure‑focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Exposure to operating or supporting ML pipelines, model‑serving infrastructure, or LLM‑based systems in production.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on‑call or production‑support experience.
Preferred Qualifications
  • Hands‑on experience operating or supporting Ray (training/serving), LangGraph or similar agent orchestration frameworks, RAG architectures, embeddings platforms, or vector store platforms.
  • Familiarity with MCP‑based tooling and ML lifecycle/dataset management systems.
  • Experience with S3 and cloud storage/networking fundamentals.
  • Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
  • Working knowledge of CI/CD pipelines and deployment workflows.
  • Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple Services Engineering (ASE) Compute - Software Engineering Manager

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 240,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Senior Software Engineer, Apple Data Platform
Senior Software Engineer, Apple Data Platform

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

JobCubby • Sunnyvale (CA), Northern (KY)

Hybrid
USD 140,000 - 230,000
Senior DevOps Engineer, Infrastructure Services Apple
Senior DevOps Engineer, Infrastructure Services Apple

Quest Technology Management • Elk Grove (CA)

Hybrid
USD 140,000 - 190,000
Site Reliability Engineer, AiDP Production Engineering
Site Reliability Engineer, AiDP Production Engineering

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000
Observability SRE Manager, Apple Services Engineering
Observability SRE Manager, Apple Services Engineering

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 260,000