Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc.

Paris

On-site

EUR 120,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Apple Inc. is seeking a Senior Machine Learning Engineer for Foundation Models Inference within Cloud OS and AI Inference. You will own end-to-end inference efficiency, hardware/software codesign, and systems architecture to deliver high-performance, privacy-preserving AI workloads at scale.

You will collaborate with Foundation Model Research and external partners, build production-grade inference systems, and mentor engineers across the organisation.

Qualifications

  • Experience leading complex ambiguous ML projects end-to-end.
  • Proficiency in PyTorch or JAX.
  • Experience working with inference frameworks.
  • Experience in Python / Go / Rust.
  • Experience deploying applications on cloud platforms (AWS, GCP or equivalent) using Kubernetes and Docker.

Responsibilities

  • Partner with Foundation Model Research to optimise inference for latest model architectures across language, vision, and speech.
  • Design and ship production-grade inference systems serving millions of customers in real time.
  • Build profiling tools and simulators to identify performance bottlenecks across hardware configurations.
  • Drive technical decisions on high-throughput, low-latency serving at scale.
  • Mentor and grow engineers across the organisation.

Skills

Python
PyTorch
JAX
LLM inference
System design
Go/Rust

Education

MS in Computer Science / ML / AI

Tools

Docker
Kubernetes
TensorRT-LLM
NVIDIA Triton
vLLM
TGI

Job description

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Paris, Ile-de-France, France Machine Learning and AI

We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant, secure and private inference stack that powers Siri AI, Apple Intelligence and Apps that are powered with the largest foundation models. Our systems serve billions of queries daily across Siri AI, Apple Intelligence, Apple Search, Apple Music, Apple TV, App Store, iMessage, Photos, Camera, Spotlight & Safari, at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimise language, vision, and speech models with billions of parameters using state-of-the-art techniques and ship them at Apple scale. This is a rare opportunity to directly shape how AI reaches billions of people worldwide.

Description

You will work at the intersection of research and production, partnering closely with the Foundation Model Research team and our external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment. You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers around you. This role sits within CloudOS and Private Cloud Compute (PCC) — Apple's purpose-built, privacy-preserving cloud infrastructure for AI workloads. PCC represents a first-of-its-kind approach to running foundation models in the cloud with verifiable privacy guarantees, and CloudOS is the systems foundation that makes it possible. You will be building and optimising inference systems on top of this infrastructure, working closely with platform and security teams to deliver both performance and trust at scale.

Responsibilities
  • Partner with the Foundation Model Research team and our external partners to optimise inference for the latest model architectures across language, vision, and speech.
  • Design and ship production-grade inference systems serving millions of customers in real time.
  • Build profiling tools and simulators to identify and resolve performance bottlenecks across different hardware configurations and use cases.
  • Drive technical decisions on high-throughput, low-latency serving at supercomputing scale.
  • Mentor and grow engineers across the organisation.
Minimum Qualifications
  • Experience leading complex ambiguous Machine learning projects end to end.
  • Proficiency in PyTorch or JAX
  • Experience working with Inference frameworks
  • Experienced in Python / Rust / Go lang or similar programming languages
  • Proficiency in deploying applications on cloud platforms (AWS, GCP or equivalent) using K8S and docker.
Preferred Qualifications
  • Hands-on experience with LLM inference stacks.
  • Working knowledge of GPU or TPU programming concepts.
  • Experience building and operating high-throughput services at large distributed scale.
  • Experience building productions systems in Go or Python.
  • Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants.
  • Experience with inference optimization frameworks such as TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server.
  • MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field

At Apple, we're not all the same. And that's our greatest strength. We draw on the differences in who we are, what we've experienced, and how we think. Because to create products that serve everyone, we believe in including everyone. Therefore, we are committed to treating all applicants fairly and equally. We will work with applicants to make any reasonable accommodations.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple • Paris

On-site
EUR 120,000 - 180,000
Senior ML Engineer - Foundation Model Inference at Scale
Senior ML Engineer - Foundation Model Inference at Scale

Apple Inc. • Paris

On-site
EUR 120,000 - 190,000
Senior ML Engineer: Foundation Model Inference at Scale
Senior ML Engineer: Foundation Model Inference at Scale

Apple • Paris

On-site
EUR 120,000 - 180,000
Internship - Machine Learning Research
Internship - Machine Learning Research

Apple Inc. • Paris

On-site
EUR 16,000 - 20,000
AIML - Machine Learning Researcher, MLR
AIML - Machine Learning Researcher, MLR

Apple • Paris

On-site
EUR 90,000 - 130,000
Offensive Security Researcher - Userland & Kernel Security
Offensive Security Researcher - Userland & Kernel Security

Apple Inc. • Paris

Hybrid
EUR 90,000 - 130,000
Offensive Security Researcher - Server Platform & Secure Computing, SEAR
Offensive Security Researcher - Server Platform & Secure Computing, SEAR

Apple Inc. • Paris

On-site
EUR 95,000 - 150,000
AI/ML Engineer - AI Systems for Security, SEAR
AI/ML Engineer - AI Systems for Security, SEAR

Lex • Paris

On-site
EUR 120,000 - 180,000
Applied Machine Learning Engineer - Security
Applied Machine Learning Engineer - Security

Apple • Paris

On-site
EUR 70,000 - 90,000
Security Tooling - Software Engineer, SEAR
Security Tooling - Software Engineer, SEAR

Apple Inc. • Paris

On-site
EUR 90,000 - 130,000