Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects

Apple Inc.

Cupertino, Northern (CA, KY)

Hybrid

USD 238,000 - 402,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical & Dental
Retirement benefits
Stock programs
Tuition reimbursement
Relocation support
Employee discounts

Job summary

Apple seeks a Senior Infrastructure, SRE & AI Platforms Manager in Cupertino, CA to shape long-term technical strategy and lead a global, mission-critical infrastructure program on the Services Special Projects team.

You will oversee compute, storage, networking, observability, and AI workloads, champion automation and reliability, and partner with leadership to align platform capabilities with business goals.

Qualifications

  • MS Degree in Computer Science or related field with 12+ years of leadership experience.
  • 6+ years managing multi-layered engineering organizations with a proven track record across global sites.
  • Hands-on and architectural mastery of cloud-native infrastructure, Kubernetes, and hybrid cloud operations.
  • Direct experience running AI training and inference workloads at scale.
  • Deep background in Site Reliability Engineering, telemetry, observability, and disaster recovery.
  • Exceptional ability to bridge executive strategy with technical trade-offs.

Responsibilities

  • Define the long-term technical vision for global compute, storage, network, observability, and AI infrastructure.
  • Lead a multi-tiered, globally distributed team and align platform capabilities with business goals.
  • Evaluate emerging infrastructure technologies and make strategic trade-off decisions.
  • Oversee AI compute at scale, scheduling, storage throughput, and capacity planning.
  • Architect hybrid/multi-cloud compute and provide seamless developer experiences.

Skills

Leadership & People Management
Kubernetes & Cloud Native
AI Infrastructure & ML workloads
SRE Principles & Telemetry
Technical Communication
Strategic Architecture

Education

MS Degree in Computer Science or related

Tools

AWS EKS
GCP GKE
Bare-metal infrastructure

Job description

Cupertino, California, United States Software and Services

We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long‑term technical strategy, organizational structure, and operational roadmap for global, mission‑critical infrastructure platforms on the Services Special Projects team.This position requires a rare blend of deep technical domain expertise—spanning distributed systems, Kubernetes, and AI workload orchestration—and proven organizational leadership managing large, globally distributed engineering teams.

Description

In this role, you will be responsible for defining and building infrastructure strategy that balances continuous innovation with high reliability, performance, and cost efficiency. You will lead a growing, multi-tiered team of engineers who are responsible for foundational platforms that power large-scale consumer and enterprise workloads.Beyond operational delivery, you will establish standards for operational excellence, Site Reliability Engineering (SRE), and capacity planning. You will be a key strategic partner, translating complex business imperatives into scalable platform designs while cultivating a strong engineering culture focused on automation, technical ownership, accountability, and continuous improvement.

Responsibilities
  • Strategic Leadership & Architecture
  • Multi-Year Roadmap & Strategy: Define and execute the long-term technical vision and capital investment strategy for global compute, storage, network, observability, and AI infrastructure.
  • Management & Organizational Alignment: Partner with Leadership to align platform capabilities, risk management, capacity investments, and architectural decisions with overarching business goals.
  • Technical Tradeoffs: Evaluate emerging infrastructure technologies, and make strategic platform trade-off decisions.
  • AI Compute & Modern Infrastructure Platforms
  • AI Infrastructure at Scale: Architect, scale, and optimize large-scale environments for training and inference, resolving complex challenges in cluster design, scheduling, interconnect performance, storage throughput, and capacity planning.
  • Hybrid & Multi-Cloud Compute: Oversee internal Kubernetes compute environments as well as managed public cloud platforms (AWS EKS, GCP GKE) and large bare-metal footprints to provide seamless developer experiences.
  • Data & Storage Platform Management: Direct the strategy and maintenance for distributed block/object storage alongside managed database and data streaming platforms (e.g., Cassandra, FoundationDB, Redis, PostgreSQL, MongoDB, Kafka).
  • Networking, Traffic & Security: Ensure reliable global traffic management, load balancing, cloud networking architectures, and enterprise security compliance across all environments.
  • SRE, Operational Excellence & Engineering Culture
  • Site Reliability Engineering (SRE): Cultivate a mature SRE culture focusing on high availability, automated fault recovery, telemetry, logging, metrics, and rigorous post-incident analysis.
  • Global Team & Leadership Development: Build, mentor, and lead a globally distributed organization comprising engineers, managers, and managers-of-managers across all levels (interns through senior principal staff).
  • Culture of Ownership & Automation: Establish an environment characterized by strong technical ownership, clear accountability, continuous operational refinement, and aggressive automation of manual processes.
Minimum Qualifications
  • MS Degree in Computer Science or related degree and 12+ years of experience of progressive engineering leadership experience building, scaling, and operating mission-critical infrastructure platforms and global services.
  • Management & Leadership Scope: 6+ years managing multi-layered engineering organizations (manager-of-managers) with a proven track record of hiring, developing, and retaining top-tier technical talent across global sites.
  • Cloud & Distributed Compute Expertise: Demonstrated hands-on and architectural mastery of cloud-native infrastructure, Kubernetes platform engineering, and hybrid cloud operations (AWS, GCP, private data centers).
  • Accelerated Computing & AI Infrastructure: Direct operational and architectural experience running large-scale systems for AI/ML training and inference workloads, including utilization optimization, scheduling, and high-performance storage/networking.
  • SRE & Production Operations: Deep background in Site Reliability Engineering (SRE) principles, telemetry, observability frameworks, disaster recovery, and managing 24/7 high-availability infrastructure at scale.
  • Technical Communication: Exceptional ability to seamlessly bridge executive strategy and low-level technical trade-offs—communicating vision to executive stakeholders while driving detailed technical discussions with principal engineers.
Preferred Qualifications
  • Large-Scale Enterprise Provenance: Experience leading core infrastructure or foundational platform SRE for a global, tier-1 technology organization operating at massive scale.
  • Multi-Engine Database & Data Infrastructure: Familiarity overseeing diverse open-source and proprietary storage/data ecosystems (e.g., Cassandra, FoundationDB, Kafka, Redis, PostgreSQL).
  • Financial & Capacity Governance: Proven competency managing large-scale infrastructure investments, capital expenditures, operational budgets, capacity forecasting, and cloud optimization strategies.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $237,600 and $401,700, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.

Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.

You’ll also receive benefits including:

  • Comprehensive medical and dental coverage
  • retirement benefits
  • a range of discounted products and free services
  • for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition
  • discretionary bonuses or commission payments as well as relocation

Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Platform Reliability Engineering, AiDP
Software Engineer, Platform Reliability Engineering, AiDP

Apple Inc. • Sunnyvale (CA)

On-site
USD 185,000 - 278,000
Medical & dental coverage
Apple stock programs
Relocation assistance
+2
Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple Services Engineering (ASE) Compute - Software Engineering Manager

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 356,000
Medical and dental coverage
Retirement benefits
Discounted products and services
+1
Software Engineering Manager, Data Solutions & Initiatives
Software Engineering Manager, Data Solutions & Initiatives

Apple Inc. • Cupertino (CA)

On-site
USD 228,000 - 343,000
Comprehensive medical and dental coverage
Employee stock purchase plan
Educational reimbursement
+1
Observability SRE Manager, Apple Services Engineering
Observability SRE Manager, Apple Services Engineering

Apple Inc. • Seattle (WA)

On-site
USD 226,000 - 338,000
Senior Software Engineer - Compute
Senior Software Engineer - Compute

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 142,000 - 263,000
Site Reliability Engineer, Customer Systems
Site Reliability Engineer, Customer Systems

Apple Inc. • Sunnyvale (CA)

On-site
USD 147,000 - 221,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+2
Senior Engineering Manager - Solutions Engineering, Apple Data Platform
Senior Engineering Manager - Solutions Engineering, Apple Data Platform

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 402,000
Medical and dental coverage
Employee stock programs
Stock Purchase Plan
+1
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

Apple Inc. • Sunnyvale (CA)

On-site
USD 216,200 - 324,800
ASE Compute - Senior SRE Software Engineer
ASE Compute - Senior SRE Software Engineer

Apple Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1