Site Reliability Engineer, AiDP Production Engineering

Apple Inc.

Austin (TX)

On-site

USD 140,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. is seeking a Site Reliability Engineer for AiDP Production Engineering in the Austin area. You will engineer resilient, scalable, real-time and batch data pipelines across bare‑metal, cloud, and Kubernetes environments to support Apple’s critical business functions.

You’ll optimize performance, manage incidents, and partner with cross‑functional teams to design robust architectures that power analytics at scale for Apple’s global operations.

Qualifications

  • 4+ years experience in cloud-native services and ETL frameworks like Apache Spark and Flink.
  • 4+ years experience in messaging systems (Kafka) and cloud infra (AWS, GCP) & Kubernetes.
  • 4+ years in modern distributed databases (Snowflake, Cassandra, SingleStore, SAP HANA).
  • 4+ years programming in Python or Java.
  • BS/MS in computer science or equivalent experience.

Responsibilities

  • Understand application requirements (Performance, Security, Scalability) and select suitable AWS/Baremetal/Kubernetes topologies.
  • Build automation for self-healing systems and infrastructure reliability.
  • Develop monitoring tools for high performance and low-latency apps.
  • Troubleshoot production issues across software, networks, and data pipelines.
  • Collaborate with teams to prioritize defects and drive incident resolution.

Skills

Python
Java
Problem solving
Communication

Education

BS/MS in Computer Science

Tools

Apache Spark
Flink
Kafka
AWS
GCP
Kubernetes
Snowflake
Cassandra
SingleStore
SAP HANA

Job description

Site Reliability Engineer, AiDP Production Engineering

Austin Metro Area, Texas, United States Software and Services

The Production Engineering team within the AI and Data Platform (AiDP) organization manages a wide array of real-time, near real-time, and batch analytical solutions. These platforms are integral to core business functions across Apple. These include sales, operations, finance, AppleCare, marketing, and services, and are instrumental in driving critical, data-driven decisions. To build these solutions, we leverage a combination of proprietary and leading open-source technologies such as Kafka, Spark, Iceberg, and Airflow. A key part of our mission is to enable AI-centric automations that enhance the overall efficiency and intelligence of the platform. We are looking for passionate engineers who thrive on solving complex infrastructure challenges at scale, both on-premises and in the cloud. If you are dedicated to optimizing scalable, maintainable, and user-friendly systems, you will find compelling opportunities to make a significant impact at AiDP.

Description

The Service Reliability Engineer (SRE) role within AiDP Production Engineering is a dynamic position that blends strategic architectural design with hands‑on technical execution. As an SRE, you will be responsible for configuring, tuning, and ensuring the resilience of complex, multi‑tiered systems to achieve optimal application performance, stability, and availability. Our team manages critical data pipelines and applications across both bare‑metal and cloud computing platforms, delivering essential data processing for all of Apple’s key business functions. We operate at an immense scale, handling exabytes of data, petabytes of memory, and tens of thousands of jobs to enable predictable and performance data analytics that power features and inform decisions across the company. If you are passionate about designing, building, and running data infrastructure that has a direct and significant impact on Apple’s global business operations, this is the ideal opportunity for you.

Responsibilities
  • Ability to understand the application requirements (Performance, Security, Scalability etc.) and assess the right services/topology on AWS, Baremetal & Kubernetes.
  • Build automation to enable self‑healing systems.
  • Build tools to monitor high performance & alert the low latency applications.
  • Ability to troubleshoot application specific, core network, system & performance issues.
  • Involvement in challenging and fast paced projects supporting Apple’s business by delivering innovative solutions.
  • Partner with engineering teams to prioritize and fix production defects.
  • Take knowledge transition from engineering teams for changes being rolled out in production.
  • Triage incidents based on the impact, devise and implement mitigation steps to unblock the business.
  • Conduct RCA, log defects and partner with engineering team for prioritization.
  • Support java based applications & Spark/Flink jobs on Baremetal, AWS & Kubernetes.
  • Share on‑call rotation with other team members to support apps and services in scope.
Minimum Qualifications
  • 4+ years experience in cloud‑native services, including ETL frameworks like Apache Spark, and Flink.
  • 4+ years experience in messaging systems (Kafka) and cloud infrastructure & services, AWS, GCP, Kubernetes.
  • 4+ years of experience in modern & distributed databases such as Snowflake, Cassandra, SingleStore, and SAP HANA.
  • 4+ years of programming experience in Python or Java.
  • BS/MS in computer science or equivalent experience.
Preferred Qualifications
  • Solid understanding of system design, data structures, and incident management best practices.
  • Should be able to understand complex architectures and be comfortable working with multiple teams.
  • Observability tools (e.g: Prometheus, Grafana, CloudWatch).
  • Ability to conduct performance analysis and troubleshoot large scale distributed systems.
  • Should be highly proactive with a keen focus on improving uptime/availability of our mission critical services.
  • Strong expertise in troubleshooting complex production issues.
  • Excellent problem solving, critical thinking, and communication skills.
  • Proven ability to resolve incidents, perform root cause analysis, and drive system reliability improvements.
  • Experience using GenAI or automation tools for issue detection, alerting, or remediation.
  • Experience in data visualization tools such as Tableau, Business Objects, ThoughtSpot.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right.

You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Platform SRE, AI & Data Platforms (AiDP)
Data Platform SRE, AI & Data Platforms (AiDP)

Apple Inc. • Austin (TX)

On-site
USD 100,000 - 130,000
Software Engineer - AiDP Reliability Engineering, IS&T, Early Career Opportunities
Software Engineer - AiDP Reliability Engineering, IS&T, Early Career Opportunities

Apple Inc. • Austin (TX)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering
Senior Site Reliability Engineer, Storage SRE / Apple Services Engineering

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)

Apple Inc. • Austin (TX)

On-site
USD 90,000 - 130,000
Site Reliability Engineer (Edge Services), Infrastructure Services
Site Reliability Engineer (Edge Services), Infrastructure Services

Apple Inc. • Elk Grove (CA)

On-site
USD 132,000 - 245,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
Software Development Engineer, AI & Data Platforms (AiDP)
Software Development Engineer, AI & Data Platforms (AiDP)

Apple Inc. • Austin (TX)

On-site
USD 90,000 - 130,000
Data & Analytics Engineer, AiDP
Data & Analytics Engineer, AiDP

Apple Inc. • Austin (TX)

On-site
USD 100,000 - 130,000
Site Reliability Engineer, Customer Systems
Site Reliability Engineer, Customer Systems

Apple Inc. • Sunnyvale (CA)

On-site
USD 147,000 - 221,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+2
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)

Apple Inc. • Sunnyvale (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and services
+1
Service Reliability Engineer (SRE)
Service Reliability Engineer (SRE)

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 264,000
Medical and Dental coverage
Retirement benefits
Employee stock purchase plan
+2