Machine Learning Platform Engineer, Apple Services Engineering

Apple Inc.

Seattle (WA)

On-site

USD 180,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. in Seattle, WA seeks an experienced software engineer to build and maintain the evaluation platform for Apple’s generative AI and agent systems.

You will own features end-to-end, write a lot of Python, and collaborate with research, model serving, and infra teams to deliver reliable services. You’ll productionize ML research, translate prototype code into scalable Python, and balance speed with careful design, tests, and rollouts.

Qualifications

  • 4-8 years of software engineering experience building and shipping production services.
  • Strong Python skills with FastAPI and Pydantic; written clean, tested code.
  • Builder mindset with quick iteration on scoped problems and care for system quality.
  • Fluency with AI coding tools and agent-oriented workflows; intuition for tool usage.
  • Familiarity with agentic LLM landscapes and production evaluation of models.
  • Hands-on experience with evaluation frameworks and trustworthy ML instrumentation.
  • Solid fundamentals: testing, CI/CD, Docker, and basic observability.
  • Clear communicator with proactive ownership and ability to surface blockers.

Responsibilities

  • Build and ship features within the evaluation platform: APIs, SDKs, orchestration components, and evaluation runners.
  • Productionize ML research by turning prototypes into reliable services.
  • Move fast with responsible decision-making, balancing speed with robust design and rollout.
  • Improve the codebase by fixing flaky tests, slow builds, and confusing APIs.
  • Develop SDKs and abstractions used by Apple teams to evaluate models and agents.
  • Ensure production readiness with tests, CI, metrics, and operational hygiene.

Skills

Python
FastAPI
Pydantic
Docker
CI/CD
Testing
Observability

Tools

Kubernetes
LangSmith
Weights & Biases

Job description

Seattle, Washington, United States Software and Services

We're building the evaluation platform that will serve all of Apple's generative AI and agent systems. Evaluating non-deterministic AI systems is one of the hardest unsolved problems in production ML — and one Apple has to get right at scale. We're building the platform that makes it tractable for every team here. This is a hands‑on engineering role with a lot of autonomy. You'll write a lot of Python and own meaningful pieces of the platform end‑to‑end. You'll be partnering closely with research engineers, model and serving teams, product and feature teams, and the infra and data platform groups this work integrates with.

Description

Build and ship: Take ownership of features and services within the evaluation platform: APIs, SDKs, orchestration components, evaluation runners. You'll have the room to make calls on your own work and the support to deliver it well. Productionize ML research: Partner with research engineers to take their prototype code and turn it into reliable services. You'll learn their world quickly and translate research patterns into clean Python that holds up under real load. Move fast, responsibly: You'll get scoped problems with room to figure out the how. We trust you to balance speed with care, to know when something needs a quick prototype and when it needs a design doc, tests, and a careful rollout. Improve as you go: Notice the rough edges and pick them up. The flaky test, the slow build, the confusing API, the runbook that's out of date. We want someone who leaves the codebase a little better every week. Developer experience: Help build the SDKs and abstractions that other Apple teams use to evaluate their models and agents. You'll feel the friction of bad ergonomics directly, which puts you in a great position to fix it. Operational ownership: Your code runs in production. You write the tests, set up the CI, add the metrics, and stay close when something breaks. You don't need to be an SRE, but you take care of what you ship.

Responsibilities
  • Build and ship: Take ownership of features and services within the evaluation platform: APIs, SDKs, orchestration components, evaluation runners. You'll have the room to make calls on your own work and the support to deliver it well.
  • Productionize ML research: Partner with research engineers to take their prototype code and turn it into reliable services. You'll learn their world quickly and translate research patterns into clean Python that holds up under real load.
  • Move fast, responsibly: You'll get scoped problems with room to figure out the how. We trust you to balance speed with care, to know when something needs a quick prototype and when it needs a design doc, tests, and a careful rollout.
  • Improve as you go: Notice the rough edges and pick them up. The flaky test, the slow build, the confusing API, the runbook that's out of date. We want someone who leaves the codebase a little better every week.
  • Developer experience: Help build the SDKs and abstractions that other Apple teams use to evaluate their models and agents. You'll feel the friction of bad ergonomics directly, which puts you in a great position to fix it.
  • Operational ownership: Your code runs in production. You write the tests, set up the CI, add the metrics, and stay close when something breaks. You don't need to be an SRE, but you take care of what you ship.
Minimum Qualifications
  • 4-8 years of software engineering experience building and shipping production services.
  • Strong Python. You're fluent with FastAPI, Pydantic, and the modern Python ecosystem. You write code that's clean, tested, and easy for the next person to pick up.
  • Builder's mindset. You enjoy shipping. You're comfortable iterating quickly on scoped problems and knowing when to slow down for the parts that need it.
  • Fluency with AI coding tools. You actively use tools like Claude Code (or equivalents) in your day‑to‑day workflow, including features like skills, slash commands, and agent‑style workflows. You have a good intuition for when to lean on them, when to steer them, and how to get high‑quality output.
  • Familiarity with the agentic LLM landscape. You stay current on how modern LLM systems work in production — tool use, MCP servers, agent frameworks, context management, multi‑step reasoning. You can hold a real conversation about the tradeoffs.
  • Hands‑on evaluation experience. You've built evaluations for your own agents or LLM systems, or you've worked with evaluation orchestration frameworks like Inspect, Braintrust, LangSmith, Promptfoo, or equivalents (including internal tooling). You understand what makes an evaluation trustworthy vs. theatrical.
  • Real working knowledge of LLMs in production. You're comfortable with prompt iteration, dataset curation, judge models, and statistical reasoning about non‑deterministic outputs. You understand the lifecycle around models even if you haven't trained them yourself.
  • Solid engineering fundamentals. You understand testing, CI/CD, containerization (Docker), and basic observability. You've shipped services that others depend on and stayed close when they broke.
  • Clear communicator. You write clear PRs, ask sharp questions, and flag blockers early. You're comfortable disagreeing thoughtfully and changing your mind when the argument is good.
  • Ownership. When something is broken or unclear, you tend to pick it up rather than wait. You either move it forward or surface it clearly.
Preferred Qualifications
  • Experience working on developer platforms, internal tools, or SDKs
  • Production experience with LLM/agent systems — building, evaluating, or operating them
  • Familiarity with job orchestration frameworks (Temporal.io, Airflow, or similar)
  • Distributed compute experience (Ray, Dask, or Kubernetes‑based job systems)
  • Experience with experiment tracking or ML lifecycle tooling (Weights & Biases, MLflow, etc.)
  • Startup or early‑stage experience where you wore multiple hats and shipped under constraint

At Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits.

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant.

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace. Learn about reasonable accommodations for job applicants.

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Platform Engineer, AI Evaluation Platform (All levels)
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 263,300
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Lead Forward Deployed Engineer, AI Evaluation Platform
Lead Forward Deployed Engineer, AI Evaluation Platform

Apple Inc. • Seattle (WA)

Hybrid
USD 175,000 - 263,300
Apple stock programs
Discretionary bonuses
Relocation assistance
Applied AI & Data Engineer - Business & Education
Applied AI & Data Engineer - Business & Education

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Apple Benefits
Relocation assistance
Discretionary bonuses
AIML - Sr Machine Learning Engineer, Evaluation
AIML - Sr Machine Learning Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 212,000 - 387,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Senior Applied Scientist - AI Evaluation & Quality Systems
Senior Applied Scientist - AI Evaluation & Quality Systems

Apple Inc. • Seattle (WA)

On-site
USD 142,300 - 263,300
AIML - Senior Machine Learning Infrastructure Engineer -ML Compute, ML Platform & Technology
AIML - Senior Machine Learning Infrastructure Engineer -ML Compute, ML Platform & Technology

Apple Inc. • Santa Clara (CA)

On-site
USD 150,400 - 277,600
AIML - Sr. Software Engineer, ML Platform Technologies (MLPT)
AIML - Sr. Software Engineer, ML Platform Technologies (MLPT)

Apple • San Francisco (CA)

On-site
USD 171,000 - 303,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and services
+1
Senior / Staff Machine Learning Engineer
Senior / Staff Machine Learning Engineer

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 309,000
Applied AI Engineer - iCloud Data
Applied AI Engineer - iCloud Data

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 175,000 - 309,000