ML Infrastructure Engineer - ML Compute Capacity

Apple

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Apple is seeking an experienced ML infrastructure engineer to scale and operate production systems for large-scale ML training and inference. You will design data pipelines, observability, and tooling across a multi-tenant fleet, collaborating with finance, data center operations, and engineering teams.

The role emphasizes capacity planning, forecasting, and cost attribution, with a focus on building self-service platforms and scalable services that stay highly available.

Qualifications

  • 7+ years of experience in machine learning infrastructure and distributed systems.
  • Experience with ML infrastructure on GPUs or TPUs.
  • Proficiency in Python and/or Go for production backend and data engineering.
  • Experience building data pipelines and queries over large-scale data (Trino, PostgreSQL, Elasticsearch).
  • Experience with observability tools (Prometheus, Grafana) or equivalents.
  • Excellent problem-framing and problem-solving skills.
  • Strong CS fundamentals.
  • Bachelor's degree or higher in Engineering, Mathematics, Economics, or related quantitative field.

Responsibilities

  • Build and operate demand and capacity planning systems.
  • Develop data pipelines and telemetry systems that ingest, normalize, and serve fleet-wide utilization and cost data.
  • Develop observability infrastructure — monitoring, alerting, and dashboards.
  • Drive forecasting, optimization, and supply chain tooling at scale.
  • Build end-to-end tooling — from data models and APIs to dashboards.
  • Build self-service platforms with well-defined schema contracts and APIs.
  • Engage cross-functionally with finance, data center operations, and infra teams.
  • Support the team through code reviews and knowledge sharing.

Skills

Python
Go
Data pipelines
Distributed systems
Observability
Problem solving
CS fundamentals

Education

Bachelor's degree or higher in Engineering, Mathematics, Economics, or related quantitative field

Tools

Trino
PostgreSQL
Elasticsearch
Prometheus
Grafana
Kubernetes
React

Job description

Summary

Scaling machine learning workloads across thousands of accelerators creates challenges that few engineers ever encounter. In Apple’s Machine Learning Platform Technologies organization, we build the infrastructure that powers large-scale ML training and inference workloads, bringing together expertise in distributed systems, machine learning infrastructure, and high-performance computing.


Description

As an engineer on the ML Compute Capacity team, you will design, build, and operate the production systems that ensure compute resources are optimally distributed throughout the company. You'll work across the stack — from data pipelines and backend services to APIs and interactive frontends — developing telemetry systems, optimization algorithms, policies, and intuitive tools for managing demand and improving efficiency across Apple's largest accelerator fleet. Our small, nimble team works in a high-autonomy, fast-paced environment, and we're passionate about digging into data patterns, laying out the performance characteristics of an entire distributed system, and knowledge sharing. If the opportunity to own and operate services that scale, stay highly available, and "just work" excites you, then please reach out to us!


Key Responsibilities


  • Build and operate demand and capacity planning systems

  • Build data pipelines and telemetry systems that ingest, normalize, and serve fleet-wide utilization and cost data across multi-tenant and heterogeneous fleets

  • Develop observability infrastructure — monitoring, alerting, and dashboards — that surfaces real-time fleet health and efficiency signals

  • Drive innovation in forecasting, optimization, and supply chain management tooling that works at scale

  • Build end-to-end tooling — from data models and APIs to interactive dashboards — that distills complex data into actionable insights for leadership

  • Build self-service platforms with well-defined schema contracts and APIs, enabling ML teams, infrastructure engineers, and finance to balance usability, utilization, and costs

  • Engage cross-functionally with finance analysts, supply chain managers, data center operations, compute infrastructure engineers, and more

  • Support the team through code reviews and knowledge sharing


Minimum Qualifications


  • 7+ years of experience in relevant areas

  • Experience with machine learning infrastructure on GPUs or TPUs

  • Proficiency in Python and/or Go for production backend and data engineering work

  • Experience building data pipelines and crafting robust queries over large-scale, multi-source data (e.g., Trino, PostgreSQL, Elasticsearch)

  • Experience with observability tools (e.g., Prometheus, Grafana) or equivalent monitoring systems

  • Excellent problem-framing and problem-solving skills

  • Strong CS fundamentals

  • Bachelor's degree or higher in Engineering, Mathematics, Economics, or a related quantitative field


Preferred Qualifications


  • Experience operating Kubernetes at production scale — including scheduling, resource management, and cluster debugging

  • Experience with modern web frameworks like React

  • Familiarity with accelerator utilization patterns across ML training and inference

  • Strong interest with capacity planning, cost attribution, or FinOps systems

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer - ML Compute Capacity
ML Infrastructure Engineer - ML Compute Capacity

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
ML Infrastructure Engineer – Capacity & Telemetry
ML Infrastructure Engineer – Capacity & Telemetry

Apple • Santa Clara (CA)

On-site
USD 180,000 - 240,000
ML Compute & Capacity Engineer
ML Compute & Capacity Engineer

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
ML Compute Efficiency Automation Engineer, Infrastructure & Planning
ML Compute Efficiency Automation Engineer, Infrastructure & Planning

Apple Inc. • Cupertino (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical coverage
Employee stock purchase program
Educational reimbursement
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
Staff ML Infrastructure Engineer
Staff ML Infrastructure Engineer

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
AIML - Senior Machine Learning Infrastructure Engineer -ML Compute, ML Platform & Technology
AIML - Senior Machine Learning Infrastructure Engineer -ML Compute, ML Platform & Technology

Apple Inc. • Santa Clara (CA)

On-site
USD 150,400 - 277,600
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Apple • Sunnyvale (CA)

On-site
USD 140,000 - 180,000
Mentorship opportunities
Innovative environment
Collaboration with cross-functional teams
On-device ML Infrastructure Engineer (Orchestration & Performance)
On-device ML Infrastructure Engineer (Orchestration & Performance)

Socket.dev • Cupertino (CA)

On-site
USD 140,000 - 200,000
Software Development Engineer, Compute Platform
Software Development Engineer, Compute Platform

Socket.dev • Seattle (WA)

On-site
USD 140,000 - 210,000