Staff ML Infrastructure Engineer

Apple

Cupertino (CA)

On-site

USD 250,000 - 320,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Apple is seeking a Staff ML Infrastructure Engineer to own the architecture of the platform powering Apple's large-scale model builds. You will define ingestion, versioning, lineage, and governance for petabyte-scale data, and design high-throughput data delivery to the GPU/TPU fleets.

You will mentor senior engineers, lead design reviews, and drive the platform roadmap for next-generation GenAI workloads, embeddings, and retrieval-augmented systems.

Qualifications

  • 10+ years in ML infrastructure or distributed data systems.
  • Experience building large-scale data or ML platforms in production.
  • Deep systems engineering with Python and a systems language (Rust preferred).
  • Familiarity with Parquet, Iceberg, Delta, Lance formats.
  • Strong collaboration and mentorship.

Responsibilities

  • Own the platform architecture behind large model builds across ingestion, versioning, lineage, and governance at petabyte scale.
  • Define data access and loading architecture to keep training compute-bound, not IO-bound.
  • Lead design reviews, mentor engineers, and align multiple teams on technical direction.
  • Drive platform roadmap for foundation models, multimodal data, and retrieval-augmented systems.

Skills

Python
Systems design
Performance engineering
Collaboration
Mentoring

Education

B.S./M.S./Ph.D. in Computer Science or Computer Engineering

Tools

Rust
C++
Go
Parquet
Iceberg
Delta
Lance
Arrow
Docker
Kubernetes

Job description

Summary

Join a team at the forefront of ML infrastructure and generative AI, where data and model workflows come together to enable the next generation of intelligent experiences on Apple products and services. We build robust systems that connect scalable data pipelines with advanced ML workflows, accelerating the development of real-world AI applications. Our work spans the full ML lifecycle, from experimentation to deployment, and you’ll play a key role in shaping how AI models are built, optimized, and scaled. We develop a platform for ML data and features that powers advanced GenAI applications. This includes embeddings (generation, evaluation, ANN search, multimodal support), AI Ops, efficient inference, and a modern feature platform designed to streamline experimentation and drive innovation. We’re looking for engineers and researchers passionate about generative models, data-centric ML, and intelligent systems across diverse real-world use cases. With the autonomy to experiment, the scale to make an impact, and the support to take ideas from prototype to production, you’ll work alongside a world-class team to build intelligent, flexible systems that make ML development faster, more reliable, and more creative.

Description

The Apple AI Platform team gives Apple's ML engineers and researchers the data systems and large-scale compute they need to build and ship models at Apple's bar for quality and privacy. Our team owns the data layer that large-scale model training depends on: ingestion, versioning, lineage, and governance on the way in, and high-throughput data loading into the training fleet on the way out. As a Staff ML Infrastructure Engineer, you will set the technical direction for that platform and own its hardest system-level problems, the architecture other engineers and teams build on.

Key Responsibilities

Own the architecture of the platform behind Apple's largest model builds: define how ingestion, immutable versioning, lineage, and governance work across structured, unstructured, and multimodal data at petabyte scale, so every model run is reproducible from a versioned dataset. Set the technical direction for high-throughput data delivery to Apple's largest GPU and TPU fleets: define the data access and loading architecture that keeps training compute-bound, not I/O-bound. Make the hard system-level and format calls that the whole platform inherits, columnar and lakehouse strategy, the dataset abstraction spanning structured and multimodal data, the shape of the SDK and core libraries, backed by design and proof, not just opinion. Drive technical direction and influence across the platform and partner teams (data, embeddings, features, research), and define the interfaces and contracts between them. Raise the technical bar across the team: mentor senior engineers, lead design reviews, and be the escalation point for the problems no one else can crack. Partner with research and product leadership to shape the platform roadmap for next-generation workloads: foundation models, multimodal data, and retrieval-augmented systems. Drive efficiency, reliability, and automation across the data plane and control plane that power Apple's ML fleet.

Minimum Qualifications

10+ years of work experience in machine learning infrastructure, distributed data systems, or a related field.10+ years of experience building and shipping large-scale data or ML infrastructure and platforms in production.Extensive experience architecting and delivering large-scale distributed data or ML infrastructure that multiple teams or products depend on in production.A track record of setting technical direction and driving it to delivery across teams, not just within a single component.Deep systems engineering: strong Python plus a systems language (Rust strongly preferred; C++ or Go acceptable), and hands-on performance engineering for I/O-bound workloads (Arrow, zero-copy, memory mapping, async I/O, high-throughput object storage).Deep familiarity with columnar and lakehouse formats (Parquet, Iceberg, Delta, or Lance) and the judgment to choose between them at scale.Strong working knowledge of the end-to-end ML workflow and how training and inference consume data, enough to architect data systems that serve them.Familiarity with modern ML and generative techniques (transformers, diffusion, retrieval-augmented generation, fine-tuning) at the level needed to design for those consumers.Demonstrated ability to design highly available, easy-to-use systems and to mentor and elevate the engineers around you.Strong collaboration and communication, with the ability to align multiple teams around a technical direction.B.S., M.S., or Ph.D. in Computer Science, Computer Engineering, or equivalent practical experience.

Preferred Qualifications

Experience defining data or ML platform architecture that was adopted across an organization.Deep experience with the data-loading and dataset-access layer of a modern ML framework (PyTorch, JAX, or TensorFlow).Distributed data-loading frameworks for ML: Ray Data, NVIDIA DALI, WebDataset, or Mosaic StreamingDataset.Experience feeding data to GPU or TPU fleets at scale and keeping them saturated.Data lineage and governance systems: DataHub, OpenLineage, Unity Catalog, or equivalent.Contributions to or operational experience with Spark, Daft, Polars, or DuckDB internals.Containerization and orchestration (Docker, Kubernetes).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff ML Infrastructure Engineer
Staff ML Infrastructure Engineer

Apple Inc. • Cupertino (CA), Northern (KY)

On-site
USD 185,000 - 325,000
Senior Machine Learning Engineer, Apple Cloud AI
Senior Machine Learning Engineer, Apple Cloud AI

Apple • Seattle (WA)

On-site
USD 150,000 - 190,000
ML Infrastructure Engineer - ML Compute Capacity
ML Infrastructure Engineer - ML Compute Capacity

Apple • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Architect for GenAI & Data Pipelines
Senior ML Platform Architect for GenAI & Data Pipelines

Apple • Cupertino (CA)

On-site
USD 250,000 - 320,000
On-Device ML Infrastructure Engineer (CoreML Runtime), Graphics, Games and Machine Learning
On-Device ML Infrastructure Engineer (CoreML Runtime), Graphics, Games and Machine Learning

Apple • Cupertino (CA)

On-site
USD 190,000 - 260,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Apple Inc. • Seattle (WA), Northern (KY)

On-site
USD 175,000 - 309,000
Systems Architect, Retail and Marcom Engineering
Systems Architect, Retail and Marcom Engineering

Apple, Inc. • Austin (TX)

On-site
USD 180,000 - 260,000
On-Device ML Quality Infrastructure Engineer, Graphics, Games & ML
On-Device ML Quality Infrastructure Engineer, Graphics, Games & ML

Apple • Cupertino (CA)

On-site
USD 170,000 - 230,000
Senior Backend Engineer - Generative AI Platform and Systems - Special Projects
Senior Backend Engineer - Generative AI Platform and Systems - Special Projects

Apple • Cupertino (CA)

On-site
USD 180,000 - 240,000
Staff ML Infra Platform Architect
Staff ML Infra Platform Architect

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000