ML Compute Efficiency Automation Engineer, Infrastructure & Planning

Apple Inc.

Cupertino (CA)

On-site

USD 181,100 - 318,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical coverage
Employee stock purchase program
Educational reimbursement

Job summary

Apple Inc. is seeking an ML Compute Efficiency Automation Engineer in Cupertino, California. In this pivotal role, you will optimize ML workflows to enhance compute efficiency and eliminate manual processes through automation. Your efforts will not only support Apple’s vast infrastructure but will also significantly impact the speed of development across teams.

The ideal candidate possesses a strong background in production software, problem-solving capabilities, and experience in AI tooling. Enjoy a competitive base salary, benefits, and opportunities for growth in a dynamic team.

Qualifications

  • 6+ years building production software and infrastructure systems.
  • Designed systems to eliminate manual work and improve efficiency.
  • Fluent in automation and AI tooling.

Responsibilities

  • Govern compute as code, ensuring accuracy and scale.
  • Identify and rectify ML inefficiencies in operations.
  • Build telemetry and anomaly detection for efficiency monitoring.

Skills

Problem-solving
Automation
Python
SQL
Collaboration

Education

BS in Computer Science or equivalent

Tools

Tableau
Looker
Grafana

Job description

ML Compute Efficiency Automation Engineer, Infrastructure & Planning

Cupertino, California, United States Machine Learning and AI

Apple’s Platform Acceleration & Compute Efficiency (PACE) is a high-leverage team operating at the intersection of our ML organizations, underlying compute infrastructure, and core platform tooling. Our mission is to empower Apple’s software engineering teams with efficient, scalable compute. By driving out operational friction and optimizing the broader machine learning ecosystem, we directly accelerate the pace of development for our Software and AIML organization. Foundation models are central to Apple's user experiences and maximizing the efficiency of our ML compute is paramount. Compute efficiency sits at the center of this role, ensuring that Apple’s models run as fast, reliably, and cost-effectively as possible. In this role you will tackle optimization challenges, from maximizing hardware utilization across GPUs, TPUs, and custom Apple Silicon, to shaping workload scheduling and capacity allocation for large model serving. We are looking for a particular kind of builder, an exceptional engineer who can think through hard problems and code their way past them, especially the ones involving scale and the slow manual work that quietly drains a high-leverage team. The ideal candidate treats every repeated process as a system waiting to be automated, every manual escalation as a system not yet built, and every prioritization request as a problem the right tooling can solve faster. The resulting data forms a foundation for the rapid, high‑quality decisions that empower Apple's technical and business leaders. This is a founding role. The majority of your time goes to AI automation, building the systems that turn manual operations into tooling that runs and corrects itself. Your remaining time will go towards hands‑on ML compute efficiency, working alongside senior ML efficiency engineers directly on the optimization problems behind the numbers. You will share ownership of PACE's governance and operations with our tools team who is actively building solutions with AI. The work a traditional operations team would grind through by hand, things like resource requests, allocation tracking, escalations, and efficiency reporting, you will turn into systems that run themselves and watch themselves. When you have done it well, the busywork is gone and PACE moves faster than its size says it should. Your challenging work will result in high development velocity and efficient compute, accelerating not only Apple, but also your career as well.

Description
  • Govern compute as code. Build the systems of record for resource requests, allocations, and utilization, accurate and at scale, so leadership can trust the numbers.
  • Hunt down ML inefficiency. Dig into inference and training workloads across GPUs, TPUs, and custom Apple Silicon, find where compute is wasted, trace it to a cause, and drive the fix.
  • Work the real optimization problems: scheduling, capacity allocation, and serving cost, alongside the engineers who own those systems.
  • Get rid of the toil. Replace the time‑sink workflows, triage, reporting, reconciliation, with systems that handle the routine and pull a person in only when judgment matters. Drive manual escalations toward zero instead of standing up a tiered on call org.
  • Make the data useful. Build the telemetry, schemas, and anomaly detection that surface efficiency and cost opportunities, then wire them into tooling that acts rather than just files a report.
  • Rebuild what breaks at scale. When a process buckles under Apple scale ML demand, re‑architect it so it grows with usage instead of headcount.
  • Make a lasting impact. Turn what you build into reusable tooling so the rest of the team benefits without coming back to you each time.
Minimum Qualifications
  • BS in Computer Science, Computer Engineering, or equivalent practical experience.
  • 6 or more years building production software, automation, tooling, or data and infrastructure systems.
  • A problem solver who builds first. You have designed things from scratch to wipe out manual work or get past a scale ceiling, and you can show us something you built.
  • Fluent with AI tooling. Coding assistants as part of how you already work, not something you read about.
  • Strong programming skills, Python or similar, or automation, pipelines, and tooling.
  • SQL, plus dashboards or data products in something like Tableau, Looker, or Grafana.
  • Experience designing data models or telemetry schemas for infrastructure, capacity, or utilization data.
  • Experience running complex systems in a large scale compute, cloud, or infrastructure environment.
  • Experience knowing where not to automate, and how to guardrail systems that act on their own.
  • Strong cross‑team collaborator who moves work forward through influence rather than authority, and is comfortable owning systems others rely on daily.
Preferred Qualifications
  • Production experience shipping automated or autonomous workflows
  • Understanding of ML training and inference infrastructure, GPU and TPU utilization, training throughput, scheduling efficiency, and foundation model serving
  • Experience building automated alerting or anomaly detection for infrastructure metrics
  • Experience with FinOps, capacity planning, cloud cost management, or IT governance
  • Knowledge of Django/Postgres
  • Love for open‑ended "go figure it out and build it" projects

At Apple, base pay is one part of our total compensation package and is determined within a range. The base pay range for this role is between $181,100 and $318,400, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Governance & Operations Lead, Infrastructure & Planning
Governance & Operations Lead, Infrastructure & Planning

Apple Inc. • Cupertino (CA)

On-site
USD 175,000 - 312,000
Discretionary bonuses
Relocation assistance
Education reimbursement
+1
Program Manager - Compute Governance Operations
Program Manager - Compute Governance Operations

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 176,000 - 312,000
Discretionary stock awards
Employee stock purchase plan
Medical and dental coverage
+2
AIML - Senior Machine Learning Infrastructure Engineer -ML Compute, ML Platform & Technology
AIML - Senior Machine Learning Infrastructure Engineer -ML Compute, ML Platform & Technology

Apple Inc. • Santa Clara (CA)

On-site
USD 150,400 - 277,600
Full Stack Software Engineer - ML Compute Capacity
Full Stack Software Engineer - ML Compute Capacity

Apple Inc. • Santa Clara (CA)

On-site
USD 184,700 - 324,800
Apple benefits
Stock programs
Relocation assistance
+1
Sr. Engineering Program Manager, ML Compute Infrastructure, Apple Services Engineering
Sr. Engineering Program Manager, ML Compute Infrastructure, Apple Services Engineering

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 312,000
Staff/Sr. ML Infrastructure Engineer, Foundation Model Compute Infra
Staff/Sr. ML Infrastructure Engineer, Foundation Model Compute Infra

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Stock programs
Relocation
Tuition reimbursement
+1
Senior / Staff Machine Learning Engineer
Senior / Staff Machine Learning Engineer

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 309,000
ML Software Engineer
ML Software Engineer

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 264,000
Employee stock purchase plan
discretionary stock awards
Medical and dental coverage
+3
AIML - Sr. Software Engineer, ML Platform Technologies (MLPT)
AIML - Sr. Software Engineer, ML Platform Technologies (MLPT)

Apple • San Francisco (CA)

On-site
USD 171,000 - 303,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and services
+1
AIML Senior Capacity Engineer - Apple Services Engineering
AIML Senior Capacity Engineer - Apple Services Engineering

Apple Inc. • Santa Clara (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Educational expense reimbursement