SRE Engineering Manager

Apple

Hyderabad

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking a Site Reliability Engineering (SRE) Manager in Hyderabad to support scalable systems for data analytics. This role requires a balance of leadership and technical skills, ensuring operational excellence across complex, distributed platforms.

The ideal candidate will bring over 10 years of SRE experience, with a strong track record in managing teams and driving system reliability improvements. Hands-on experience with cloud infrastructures like AWS and GCP is essential, along with skills in programming languages such as Python or Java.

Qualifications

  • 10+ years of experience in the SRE domain.
  • At least 3 years in a management role.
  • Experience with large scale distributed systems.

Responsibilities

  • Drive operational excellence and contribute to code.
  • Manage Infrastructure as Code (IaC).
  • Participate in on-call rotations and resolve critical issues.
  • Perform root cause investigations.
  • Collaborate with engineering teams.

Skills

Technical leadership
Cloud infrastructure knowledge
Incident response leadership
Programming in Python/Java/Scala

Education

Bachelor’s degree or equivalent

Tools

AWS
GCP
Kubernetes
Prometheus
Grafana

Job description

Summary

Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, or Apple Store experience we deliver is the result of us making each other’s ideas stronger. That happens because every one of us shares a belief that we can make something wonderful and share it with the world, changing lives for the better. It’s the diversity of our people and their thinking that inspires the innovation that runs through everything we do. When we bring everybody in, we can do the best work of our lives. Here, you’ll do more than join something — you’ll add something. Apple’s Artificial Intelligence and Data Platforms (AiDP) team is seeking an experienced Site Reliability Engineering (SRE) Manager to support scalable and resilient distributed systems that power Apple’s data pipelines and analytics platforms. Our Enterprise Data Warehouse landscape caters to a wide variety of real‑time, near real‑time and batch analytical solutions. These solutions are an integral part of business functions like Sales, Operations, Finance, AppleCare, Marketing and Internet Services, enabling business drivers to make critical decisions. We utilizes proprietary and open source technologies such as Kafka, Spark, Iceberg, Airflow, and others to build these solutions. If you are passionate about addressing infrastructure challenges at scale, both on‑premises and in the cloud, and focused on optimizing scalable solutions by prioritizing ease of use and maintenance, you will discover exciting opportunities in AiDP.

Description

As a hands‑on SRE Manager, you’ll lead by example—actively driving operational excellence, contributing to code, and ensuring system reliability. You will be deeply involved in incident response across complex, distributed data platforms designed to support data exploration, analytics, and reporting solutions. These platforms operate at the unique intersection of high data volume and hybrid infrastructure, spanning both cloud and on‑premise environments. We are looking for a collaborative and innovative leader who thrives under tight deadlines, excels at solving complex problems, and consistently delivers high‑quality, forward‑thinking solutions.

Responsibilities
  • Lead by Example: Provide technical leadership and guidance to SRE team by applying hands‑on skills and continuous learning. Build and mentor a world‑class engineering team that partners closely with platform teams to design scalable, reliable systems, while contributing actively to both platform and application code.
  • Drive Automation for Data Platforms and Infrastructure: Manage Infrastructure as Code (IaC) and develop tooling to enhance engineering productivity. Lead initiatives for cost optimization and operational efficiency at scale.
  • Incident Response and On‑Call Engagement: Actively participate in on‑call rotations and resolve critical production issues. Lead response efforts during major incidents and serve as the primary escalation point for complex problems.
  • Drive Post‑Incident Analysis: Perform root cause investigations and ensure follow‑up with actionable postmortems and infrastructure hardening initiatives. Implement fixes—in code, infrastructure, or processes—to prevent recurrence.
  • Active Collaboration with Cross‑Functional Teams: Partner closely with engineering teams to troubleshoot issues, deploy fixes, and enhance system reliability. Champion operational excellence through direct technical contributions.
  • Establish Production Readiness Standards: Take ownership of Application Security, Disaster Recovery & Application Documentation to reflect latest system architecture and configurations.
Minimum Qualifications
  • Bachelor’s degree or equivalent, with 10+ years of experience in the SRE domain and at least 3 years in a management role focused on leading, hiring, developing and building teams.
  • Hands‑on experience building, supporting/maintaining applications. large scale distributed systems in cloud or hybrid environments.
  • Strong knowledge of cloud infrastructure & services (e.g., AWS, GCP, Kubernetes), Observability tools (e.g: Prometheus, Grafana, CloudWatch).
  • Strong Programming experience in one of the programming languages - Python or Java or Scala.
  • Proven ability to lead incident response, perform root cause analysis, and drive system reliability improvements.
  • Able to lead across organizational boundaries and diverse reporting structures.
Preferred Qualifications
  • Hands‑on experience supporting enterprise data systems on distributed architectures.
  • Expertise in cloud‑native services, including ETL frameworks (Apache Spark, Flink), and messaging systems (Kafka).
  • Exposure to data visualization tools such as Tableau, Business Objects, ThoughtSpot, with experience supporting and troubleshooting issues related to dashboards and reports.
  • Experience with modern & distributed databases such as Snowflake, Cassandra, SingleStore, and SAP HANA.
  • Experience using GenAI or automation tools for issue detection, alerting, or remediation.
  • Solid understanding of system design, data structures, and incident management best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Manager
Site Reliability Engineering Manager

Apple • Bengaluru

On-site
INR 2,000,000 - 2,500,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Apple • Hyderabad

On-site
INR 2,500,000 - 3,500,000
Site Reliability Engineering Manager, Apple Data Platform
Site Reliability Engineering Manager, Apple Data Platform

Apple Inc. • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Site Reliability Engineer - Insights
Site Reliability Engineer - Insights

Apple • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer - AI & Data Platforms
Site Reliability Engineer - AI & Data Platforms

Apple • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Full Stack Software Engineer - Manufacturing Systems & Infrastructure
Full Stack Software Engineer - Manufacturing Systems & Infrastructure

Apple • Bengaluru

On-site
INR 1,800,000 - 2,500,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Software Engineer - AI and Data Platforms
Software Engineer - AI and Data Platforms

Apple • Hyderabad

On-site
INR 6,704,000 - 9,100,000
Accessibility-focused workplace
Comprehensive benefits
Senior Lead Software Engineer - SRE
Senior Lead Software Engineer - SRE

United States Digital Space LLC • Karnataka

On-site
INR 2,000,000 - 5,000,000