Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Apple is seeking an experienced ML infrastructure engineer to scale and operate production systems for large-scale ML training and inference. You will design data pipelines, observability, and tooling across a multi-tenant fleet, collaborating with finance, data center operations, and engineering teams.
The role emphasizes capacity planning, forecasting, and cost attribution, with a focus on building self-service platforms and scalable services that stay highly available.
Scaling machine learning workloads across thousands of accelerators creates challenges that few engineers ever encounter. In Apple’s Machine Learning Platform Technologies organization, we build the infrastructure that powers large-scale ML training and inference workloads, bringing together expertise in distributed systems, machine learning infrastructure, and high-performance computing.
As an engineer on the ML Compute Capacity team, you will design, build, and operate the production systems that ensure compute resources are optimally distributed throughout the company. You'll work across the stack — from data pipelines and backend services to APIs and interactive frontends — developing telemetry systems, optimization algorithms, policies, and intuitive tools for managing demand and improving efficiency across Apple's largest accelerator fleet. Our small, nimble team works in a high-autonomy, fast-paced environment, and we're passionate about digging into data patterns, laying out the performance characteristics of an entire distributed system, and knowledge sharing. If the opportunity to own and operate services that scale, stay highly available, and "just work" excites you, then please reach out to us!