ML Infrastructure Engineer – Capacity & Telemetry

Apple

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Apple is seeking an experienced ML infrastructure engineer to scale and operate production systems for large-scale ML training and inference. You will design data pipelines, observability, and tooling across a multi-tenant fleet, collaborating with finance, data center operations, and engineering teams.

The role emphasizes capacity planning, forecasting, and cost attribution, with a focus on building self-service platforms and scalable services that stay highly available.

Qualifications

  • 7+ years of experience in machine learning infrastructure and distributed systems.
  • Experience with ML infrastructure on GPUs or TPUs.
  • Proficiency in Python and/or Go for production backend and data engineering.
  • Experience building data pipelines and queries over large-scale data (Trino, PostgreSQL, Elasticsearch).
  • Experience with observability tools (Prometheus, Grafana) or equivalents.
  • Excellent problem-framing and problem-solving skills.
  • Strong CS fundamentals.
  • Bachelor's degree or higher in Engineering, Mathematics, Economics, or related quantitative field.

Responsibilities

  • Build and operate demand and capacity planning systems.
  • Develop data pipelines and telemetry systems that ingest, normalize, and serve fleet-wide utilization and cost data.
  • Develop observability infrastructure — monitoring, alerting, and dashboards.
  • Drive forecasting, optimization, and supply chain tooling at scale.
  • Build end-to-end tooling — from data models and APIs to dashboards.
  • Build self-service platforms with well-defined schema contracts and APIs.
  • Engage cross-functionally with finance, data center operations, and infra teams.
  • Support the team through code reviews and knowledge sharing.

Skills

Python
Go
Data pipelines
Distributed systems
Observability
Problem solving
CS fundamentals

Education

Bachelor's degree or higher in Engineering, Mathematics, Economics, or related quantitative field

Tools

Trino
PostgreSQL
Elasticsearch
Prometheus
Grafana
Kubernetes
React

Job description

Apple is seeking an experienced ML infrastructure engineer to scale and operate production systems for large-scale ML training and inference. You will design data pipelines, observability, and tooling across a multi-tenant fleet, collaborating with finance, data center operations, and engineering teams.

The role emphasizes capacity planning, forecasting, and cost attribution, with a focus on building self-service platforms and scalable services that stay highly available.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Compute & Capacity Engineer
ML Compute & Capacity Engineer

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
ML Infrastructure Engineer - ML Compute Capacity
ML Infrastructure Engineer - ML Compute Capacity

Apple • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Staff ML Infra Platform Architect
Staff ML Infra Platform Architect

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
ML Infrastructure Engineer - ML Compute Capacity
ML Infrastructure Engineer - ML Compute Capacity

Apple Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Cloud Infrastructure Engineer for ML/GenAI Platform
Cloud Infrastructure Engineer for ML/GenAI Platform

Apple Inc. • Sunnyvale (CA)

On-site
USD 185,000 - 325,000
Medical & dental coverage
Retirement benefits
Discounted products and services
+1
Senior ML Infrastructure & QA Engineer
Senior ML Infrastructure & QA Engineer

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 308,500
Employee stock purchase plan
Apple shareholder programs
Relocation assistance
Senior ML Infra Platform Engineer — Scalable AI Compute
Senior ML Infra Platform Engineer — Scalable AI Compute

Apple Inc. • Santa Clara (CA)

On-site
USD 150,400 - 277,600
Senior MLOps Engineer: Scalable Production ML Pipelines
Senior MLOps Engineer: Scalable Production ML Pipelines

Apple Inc. • Cupertino (CA)

On-site
USD 216,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
Staff ML Compute & TPU Infrastructure Engineer
Staff ML Compute & TPU Infrastructure Engineer

Apple Inc. • San Francisco (CA)

On-site
USD 210,000 - 300,000
Engineering Manager, ML Production & AIOps Platforms
Engineering Manager, ML Production & AIOps Platforms

Apple Inc. • Cupertino (CA)

On-site
USD 268,000 - 402,000
Apple stock program
Relocation assistance
Comprehensive health benefits
+1