Apple Services Engineering (ASE) Compute - Software Engineering Manager

Socket.dev

Cupertino (CA)

On-site

USD 190,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking an experienced Software Engineering Manager for its ASE Compute team. You will lead Infrastructure and SRE engineers responsible for operating and scaling large-scale compute infrastructure across data centers.

Responsibilities include capacity planning, incident management, release engineering, observability, and modernization. You will balance reliability with innovation and champion AI-driven automation across geographies.

Qualifications

  • 5+ years leading infrastructure or platform engineering teams.
  • Experience building on-call/org incident management processes.
  • Strong cloud infra, compute orchestration, and scale understanding.

Responsibilities

  • Manage and grow a distributed-infrastructure team across geographies.
  • Operate batch compute and core controllers with high availability.
  • Drive capacity planning, incident management, and release engineering.
  • Advance observability, automation, and tooling to scale platforms.

Skills

SRE leadership
Incident management
Cross-team collaboration
Talent development
Cloud infrastructure

Tools

Kubernetes
OpenStack
KVM
Chef
Ansible
Terraform
Salt
Go
Python

Job description

People at Apple don't just build products — they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!

Description

Apple Service Engineering (ASE)'s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies — ensuring platform availability, reliability, and performance at Apple scale. You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains.

Minimum Qualifications
  • 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
  • Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes
  • Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale
  • Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt)
  • Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering
  • Excellent verbal and written communication skills with the ability to influence across teams and levels
  • Demonstrated ability to recruit, develop, and retain high-performing engineering talent
Preferred Qualifications
  • Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation
  • Experience managing or scaling batch compute, job scheduling, or HPC platforms
  • Proficiency in Go or Python with a strong automation-first mindset
  • Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale
  • Experience operating large-scale multi-tenant Infrastructure as a Managed Service
  • Experience managing geographically distributed teams and follow-the-sun on-call models
  • Track record of driving capacity efficiency initiatives resulting in measurable cost optimization
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple Services Engineering (ASE) Compute - Software Engineering Manager

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 356,000
Medical and dental coverage
Retirement benefits
Discounted products and services
+1
ASE Compute - Senior SRE Software Engineer
ASE Compute - Senior SRE Software Engineer

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior/Staff Software Engineer, Apple Services Engineering
Senior/Staff Software Engineer, Apple Services Engineering

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Socket.dev • Austin (TX)

On-site
USD 140,000 - 220,000
Site Reliability Engineer, Apple Data Platform / Big Data Platform
Site Reliability Engineer, Apple Data Platform / Big Data Platform

Socket.dev • Austin (TX)

On-site
USD 120,000 - 180,000
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform

Apple • Austin (TX)

On-site
USD 120,000 - 180,000
Senior Software Engineer, Apple Services Engineering - Kubernetes
Senior Software Engineer, Apple Services Engineering - Kubernetes

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
Infrastructure & SRE Manager: Scale Global Compute
Infrastructure & SRE Manager: Scale Global Compute

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 240,000
Software Engineer, Cloud Services ASE
Software Engineer, Cloud Services ASE

Socket.dev • Austin (TX)

On-site
USD 180,000 - 240,000
AI-Driven Compute Infra Engineering Manager
AI-Driven Compute Infra Engineering Manager

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 356,000
Medical and dental coverage
Retirement benefits
Discounted products and services
+1