Software Site Reliability Engineer - AI & Data Platforms

Apple Inc.

Cupertino (CA)

On-site

USD 150,000 - 190,000

Full time

17 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Apple Inc. is seeking a Software Site Reliability Engineer for AI & Data Platforms (AiDP) to design, build, and operate a scalable data platform supporting analytics, reporting, and AI/ML apps.

You will optimize performance, automate operations, and resolve complex production issues at scale. The role requires 3+ years in SRE or software dev, proficiency in Python, Golang, or Java, and hands-on experience with distributed systems, cloud platforms, and Kubernetes-based environments.

Qualifications

  • 3+ years in software site reliability engineering or software development.
  • Proficiency in Python, Golang, or Java.
  • Experience with distributed data pipelines and cloud platforms.
  • Experience with Open Source projects and multi-tenant Kubernetes environments is a plus.

Responsibilities

  • Design, develop, and automate tools and frameworks to improve reliability and scalability of large-scale data platforms.
  • Monitor and maintain on-prem and cloud workloads with advanced monitoring and alerting.
  • Troubleshoot production incidents, perform root cause analysis, and resolve issues across analytics and AI/ML apps.
  • Collaborate with development and operations teams to integrate reliability best practices throughout the software lifecycle.
  • Proactively optimize architecture, deployment, and operations for distributed systems.

Skills

Python
Golang
Java
Distributed systems

Tools

Kubernetes
Airflow
Spark

Job description

Software Site Reliability Engineer - AI & Data Platforms

Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, is the result of us making each other’s ideas stronger. That happens because every one of us shares a belief that we can make something wonderful and share it with the world, changing lives for the better. It’s the diversity of our people and their thinking that inspires the innovation that runs through everything we do. When we bring everybody in, we can do the best work of our lives. Here, you’ll do more than join something — you’ll add something. Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless. Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here. AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

Description

As a Data Platform SRE, you will be responsible for developing and operating our big data platform using open source or other solutions to aid critical applications, such as analytics, reporting, and AI/ML apps. This includes working to optimize performance and cost, automate operations, and identifying and resolving production issues to ensure the best data platform experience

Responsibilities
  • Design, develop, and automate: Build tools, frameworks and solutions to improve reliability, scalability, and efficiency across large scale distributed data platform systems.
  • Monitor and maintain: Implement advanced monitoring and alerting for on-prem , cloud and workloads.
  • Troubleshoot and solve: Support critical applications including analytics, reporting, and AI/ML apps. Respond to and resolve complex production incidents, and perform root cause analysis.
  • Collaborate: Work closely with development and operations teams to integrate reliability best practices throughout the software lifecycle.
  • Optimize: Proactively recommend improvements in architecture, deployment, and operations for distributed systems
Minimum Qualifications
  • Experience: 3+ years in software site reliability engineering or software development roles.
  • Programming: Proficient in at least one of Python, Golang, or Java.
  • Skilled at coding for distributed systems and developing resilient data pipelines.
  • Cloud Platforms: Hands-on experience with at least one major cloud platform (AWS, Azure, or Google Cloud Platform).
Preferred Qualifications
  • Expertise in designing, building, and operating critical, large-scale distributed systems with a focus on low latency, fault-tolerance, and high availability.
  • Experience with contribution to Open Source projects is a plus.
  • Experience with multiple public cloud infrastructure, managing multi-tenant Kubernetes clusters at scale and debugging Kubernetes/Spark issues.
  • Experience with workflow and data pipeline orchestration tools (e.g., Airflow, DBT).
  • Understanding of data modeling and data warehousing concepts.
  • Familiarity with the AI/ML stack, including GPUs, MLFlow, or Large Language Models (LLMs).
  • Data Structures & Algorithms: Strong foundation and application experience.
  • Distributed Systems: Solid understanding and hands-on experience managing at least one distributed system (e.g. Kafka, Spark, Flink etc. ).
  • Solid understanding of software engineering best practices, including the full development lifecycle, secure coding, and experience building reusable frameworks or libraries.
  • Problem Solving: Demonstrated ability to independently troubleshoot and resolve complex technical issues.
  • Creative Thinking: A track record of proposing and implementing innovative solutions to technical challenges.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer : Data & AI
Software Engineer : Data & AI

Apple Inc. • Cupertino (CA)

On-site
USD 180,000 - 250,000
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)
Software Engineer (Data Solutions), AI & Data Platforms (AiDP)

Apple Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 190,000
Software Engineer, Platform Reliability Engineering, AiDP
Software Engineer, Platform Reliability Engineering, AiDP

Apple Inc. • Sunnyvale (CA)

On-site
USD 185,000 - 278,000
Medical & dental coverage
Apple stock programs
Relocation assistance
+2
Software Engineer (Framework Solutions), AI & Data Platforms (AiDP)
Software Engineer (Framework Solutions), AI & Data Platforms (AiDP)

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 170,000
Software Engineer - Data Solutions, AI & Data Platform (AiDP)
Software Engineer - Data Solutions, AI & Data Platform (AiDP)

Apple Inc. • Sunnyvale (CA), Northern (KY)

On-site
USD 150,000 - 225,000
Comprehensive medical and dental
Retirement benefits
Employee stock purchase plan
+1
Site Reliability Engineer - Insights
Site Reliability Engineer - Insights

Apple Inc. • Cupertino (CA)

On-site
USD 120,000 - 180,000
Sr. Software Engineer (Data Solutions), IS&T Ai & Data Platforms
Sr. Software Engineer (Data Solutions), IS&T Ai & Data Platforms

Apple Inc. • Sunnyvale (CA), Northern (KY)

On-site
USD 185,000 - 278,000
Site Reliability Engineer, Apple Data Platform
Site Reliability Engineer, Apple Data Platform

Apple Inc. • Austin (TX), Northern (KY)

Hybrid
USD 150,000 - 190,000
Software Engineer, Reliability Engineering, AiDP
Software Engineer, Reliability Engineering, AiDP

Apple Inc. • Sunnyvale (CA), Northern (KY)

On-site
USD 150,000 - 225,000
Stock programs
Relocation assistance
Education reimbursement
+1