SRE - AI Data Platform & ML Infra at Scale

Apple Inc.

Austin, Northern (TX, KY)

Hybrid

USD 140,000 - 180,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Data Platform SRE team is hiring in Austin to operate and support a multi-cloud data/AI platform powering internal services and ML workflows. You’ll own reliability for data pipelines, ML/AI services, and agent platforms, including Ray and LangGraph deployments, with strong emphasis on incident response and automation.

You will collaborate with cross-functional teams to design monitoring and dashboards, drive reliability roadmaps, and grow into an SME for next-gen Apple AI infrastructure.

Qualifications

  • Bachelor's Degree in Computer Science, engineering, or equivalent experience.
  • 1–4 years in Site Reliability Engineering, DevOps, or infra-focused roles.
  • Proficient in Python; Golang a plus.
  • Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
  • Exposure to ML pipelines, model-serving infra, or LLM-based systems in production.
  • Strong communication and composure during incidents.
  • Solid grounding in SRE principles with on-call/production-support experience.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio.
  • Participate in rotating on-call across supported services.
  • Own the operational health of ML/AI platform services as SME (Ray, LangGraph, RAG pipelines).
  • Provide Slack-based support to internal customers; triage and resolve issues.
  • Onboard new services with monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Build automation and self-healing tooling to reduce toil and scale capacity.
  • Identify, escalate, and resolve production issues to protect reliability.
  • Collaborate with SRE and dev teams to align execution with goals.

Skills

Python
Golang
Kubernetes
AWS
GCP
ML pipelines
LLM-based systems
On-call
Communication

Education

Bachelor's Degree in Computer Science or engineering

Tools

Prometheus
Grafana
Splunk
CI/CD
Ray
LangGraph

Job description

Apple Data Platform SRE team is hiring in Austin to operate and support a multi-cloud data/AI platform powering internal services and ML workflows. You’ll own reliability for data pipelines, ML/AI services, and agent platforms, including Ray and LangGraph deployments, with strong emphasis on incident response and automation.

You will collaborate with cross-functional teams to design monitoring and dashboards, drive reliability roadmaps, and grow into an SME for next-gen Apple AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE – AI Data Platform Infra & ML Pipelines
SRE – AI Data Platform Infra & ML Pipelines

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
SRE: Data Platform & Multi-Cloud Infra
SRE: Data Platform & Multi-Cloud Infra

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 140,000 - 180,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
Global AI Infrastructure & SRE Platform Lead
Global AI Infrastructure & SRE Platform Lead

Apple Inc. • Cupertino (CA)

On-site
USD 238,000 - 402,000
Stock options
Relocation
Site Reliability Engineer - AI Data Platform Pipelines
Site Reliability Engineer - AI Data Platform Pipelines

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000
Platform Reliability Engineer: ML & GenAI Systems
Platform Reliability Engineer: ML & GenAI Systems

Apple Inc. • Austin (TX)

On-site
USD 120,000 - 180,000
Senior SRE: AI-Driven Compute Platform
Senior SRE: AI-Driven Compute Platform

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 142,000 - 263,000
Site Reliability Engineer - Big Data Platform & Multi-Cloud
Site Reliability Engineer - Big Data Platform & Multi-Cloud

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000
Site Reliability Engineer, AiDP Production Engineering
Site Reliability Engineer, AiDP Production Engineering

Apple Inc. • Austin (TX)

On-site
USD 140,000 - 170,000