ML Infrastructure Service Reliability Engineer- Apple Services Engineering

Apple Inc.

Bengaluru

On-site

INR 2,800,000 - 4,000,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Apple Inc. in Bengaluru is seeking an experienced ML Infrastructure Site Reliability Engineer to join the Services Engineering team.

You will manage a large ML compute platform, multi-cloud storage, and a caching layer, ensuring high availability and efficiency at scale across Apple’s ecosystem. The role demands strong Kubernetes expertise, cloud storage experience (S3/GCS), and programming skills in Python, Go, or Rust.

Qualifications

  • 4+ years experience in building, operating and scaling a large application in a private, public or hybrid cloud environment
  • Deep expertise in Kubernetes, with hands-on experience using GKE or EKS
  • Proficient in Python, Go, or Rust
  • Practical experience with object storage technologies, including S3 or GCS
  • Strong background in designing and troubleshooting complex networking issues in cloud infra
  • Solid understanding of Linux internals and distributed systems

Responsibilities

  • Participates in a rotating on-call schedule, including occasional weekend coverage when necessary
  • Currently headquartered in Cupertino, with active expansion in Bangalore to support global operations across time zones
  • Leverages a diverse stack including open-source tools, commercial solutions, and internally developed systems
  • Encourages open dialogue, values strong ideas, and recognizes impactful results

Skills

Kubernetes
Python
Go
Rust
Linux
Networking

Tools

Google Kubernetes Engine (GKE)
Amazon EKS
Amazon S3
Google Cloud Storage (GCS)
Spinnaker
Helm
Flux

Job description

ML Infrastructure Service Reliability Engineer- Apple Services Engineering

At Apple, we don’t just build products — we create transformative experiences that have reshaped entire industries. Our innovation is driven by the diversity of our people and their ideas, inspiring everything we do. Imagine the impact you could make. Join Apple and help us leave the world better than we found it.The ML Infrastructure team is responsible for managing Apple’s largest ML compute platform, multi-cloud storage abstraction and caching platform, which supports critical machine learning training workloads that power user-facing features across the Apple ecosystem. Operating across both first-party and third-party cloud environments brings complex and unique challenges.As a Site Reliability Engineer (SRE) on the ML Infrastructure team, you’ll be expected to address these challenges through a strong foundation in cloud object storage, data analysis, automation, collaboration, and advanced expertise in Kubernetes. Our team oversees the full infrastructure stack — from low-level nodes to the complete network architecture — ensuring our platform remains highly available, resilient, and efficient at scale.

Description

We are seeking an experienced Software and Systems Engineer to join our dynamic team. This role demands a proactive mindset, technical excellence, and a collaborative spirit.The ideal candidate will demonstrate:Strong critical thinking and a high degree of individual accountabilityEffective communication and collaboration skillsA genuine passion for Infrastructure as a Service (IaaS)A commitment to automation and operational efficiencyOwnership of projects from design through deliveryA solutions-oriented approach, coupled with the ability to gain alignment ontechnical directionConsistent and timely execution of design implementations aligned withproject objectivesThe ability to provide constructive technical feedback, fostering team-widegrowth and continuous improvement

Responsibilities
  • Participates in a rotating on-call schedule, including occasional weekend coverage when necessary
  • Currently headquartered in Cupertino, with active expansion in Bangalore to support global operations across time zones
  • Leverages a diverse stack including open-source tools, commercial solutions, and internally developed systems
  • Encourages open dialogue, values strong ideas, and recognizes impactful results
Minimum Qualifications
  • 4+ years experience in building, operating and scaling a large application in a private, public or hybrid cloud environment
  • Deep expertise in Kubernetes, with hands-on experience using platforms such as Google Kubernetes Engine (GKE) or Amazon Elastic Kubernetes Service (EKS)
  • Proficient in designing, developing, and releasing code in languages such as Python, Go, or Rust
  • Practical experience with object storage technologies, including Amazon S3 or Google Cloud Storage (GCS)
  • Strong background in designing and troubleshooting complex networking issues in both public and private cloud infrastructures
  • Solid understanding of Linux internals, standard networking protocols, and distributed systems architecture
Preferred Qualifications
  • Proven drive to automate manual operations and enhance processes through continuous iteration
  • Strong understanding of best practices for deploying large-scale, distributed applications
  • Hands-on experience managing diverse system environments using configuration management tools or software delivery platforms such as Spinnaker, Helm, or Flux
  • Demonstrated expertise in deploying, supporting, and monitoring both new and existing services, platforms, and application stacks
  • Solid familiarity with container orchestration and management using Kubernetes

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infrastructure Service Reliability Engineer- Apple Services Engineering
ML Infrastructure Service Reliability Engineer- Apple Services Engineering

Apple • Bengaluru

On-site
INR 1,800,000 - 2,500,000
Site Reliability Engineer - Apple Services Engineering
Site Reliability Engineer - Apple Services Engineering

Apple Inc. • Bengaluru

On-site
INR 4,500,000 - 6,500,000
Database Infrastructure and Reliability Engineer- Apple Services Engineering
Database Infrastructure and Reliability Engineer- Apple Services Engineering

Apple Inc. • India

On-site
INR 4,200,000 - 6,800,000
Database Infrastructure and Reliability Engineer- Apple Services Engineering
Database Infrastructure and Reliability Engineer- Apple Services Engineering

Apple Inc. • Bengaluru

On-site
INR 1,800,000 - 2,600,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Apple Inc. • India

On-site
INR 4,000,000 - 6,000,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Apple Inc. • Hyderabad

On-site
INR 3,500,000 - 5,200,000
Software Site Reliability Engineer - AI & Data Platforms
Software Site Reliability Engineer - AI & Data Platforms

Apple Inc. • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Site Reliability Engineer - AI & Data Platforms
Site Reliability Engineer - AI & Data Platforms

Apple • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Software Site Reliability Engineer - AI & Data Platforms
Software Site Reliability Engineer - AI & Data Platforms

Apple • Bengaluru

On-site
INR 1,800,000 - 2,400,000
DevOps Engineer, Retail Engineering
DevOps Engineer, Retail Engineering

Apple • Hyderabad

On-site
INR 2,000,000 - 4,000,000