Stand out for this role — generate a tailored resume and cover letter in about a minute.
Apple is seeking a senior ML systems engineer to drive optimization for large-scale foundation model training on TPUs. You will profile and optimize JAX/XLA workloads across compute, memory, and communication, and develop high-performance TPU kernels for core ML operations.
The role requires 6+ years in high-performance ML or distributed systems, strong Python skills, and expertise in distributed systems and performance optimization.
Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build, service we create, or Apple Store experience we deliver is the result of us making each other’s ideas stronger. That happens because every one of us shares a belief that we can make something wonderful and share it with the world, changing lives for the better. It’s the diversity of our people and their thinking that inspires the innovation that runs through everything we do. When we bring everybody in, we can do the best work of our lives. Here, you’ll do more than join something — you’ll add something!
6+ years of experience building or optimizing high-performance ML or distributed systemsProficient in Python or other relevant programming languagesStrong understanding of distributed systems, parallel computing, and performance optimizationExperience profiling and optimizing compute-, memory-, or communication-intensive workloadsAbility to clearly communicate complex technical problems and collaborate with partners to develop solutionsBachelor's degree in Computer Science, Engineering, or a related field
Advanced degree in Computer Science, Engineering, or a related fieldExperience with accelerators such as TPU or GPU and understanding of accelerator architecture and performance characteristicsExperience with JAX, XLA, PyTorch or other ML compiler/runtime stacksExperience developing or optimizing accelerator kernels using Pallas, Triton, CUDA, or similar technologiesExperience optimizing large-scale foundation model training and distributed communication