AIML - Sr Machine Learning Engineering Manager, Evaluation

Apple Inc.

Cupertino, Northern (CA, KY)

Hybrid

USD 238,000 - 402,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. seeks a senior hands-on Machine Learning Engineering Manager to lead a small team focused on evaluating foundation models and agentic systems, diagnosing failure modes, and driving automated improvements to prompts, contexts, tools, and data generation.

You will translate research into scalable pipelines, mentor engineers, and align teams around a clear technical direction across Apple Foundation Models and product groups, advancing synthetic data generation and privacy-conscious

Qualifications

  • 8+ years of professional experience in machine learning, applied research, or software engineering, including production ML systems or large-scale experimentation platforms.
  • 3+ years of technical leadership, including direct people management of ML or software engineers and mentoring/talent growth.
  • Master’s or PhD in Computer Science, ML, AI, or related field.
  • Strong hands-on programming in Python and building reliable ML pipelines.
  • Deep experience with large language models or agentic systems (evaluation, planning, or action-taking workflows).
  • Experience building automated evaluation methods (LLM-based judges, rubrics, reward models, or scalable benchmarks).
  • Experience with model/agent refinement areas (prompt/context optimization, post-training, RL, or agent-harness optimization).
  • Excellent communication and collaboration across research, engineering, and product teams.

Responsibilities

  • Architect scalable evaluation systems for foundation models and agents, including benchmarks, evaluators, environments, and regression testing.
  • Establish an end-to-end evaluation flywheel connecting failures to diagnosis, refinement, and measurable quality improvements.
  • Lead and grow a small ML engineering team, while contributing to design, experimentation, implementation, and reviews.
  • Define the strategy and roadmap for automatic prompt, context, tool, rubric, and agent-harness optimization for agentic development and model evaluation.
  • Develop methods to convert evaluation findings into actionable model-improvement signals and data generation pipelines.
  • Collaborate across AIML to design scalable synthetic data generation for evaluation and post-training.
  • Apply recent research to production-quality workflows (LLM evaluation, reward modeling, test-time search, post-training).

Skills

ML leadership
Python
Production ML systems
Large language models
Experimentation platforms
Mentoring engineers
Cross-functional collaboration
Architecture & design

Education

Master’s or PhD in Computer Science, Machine Learning, AI, or related field

Tools

TensorFlow
PyTorch

Job description

AIML - Sr Machine Learning Engineering Manager, Evaluation

Cupertino, California, United States Machine Learning and AI

Apple's AIML Evaluation team builds the systems and methodologies that measure and improve the quality of foundation models and agentic experiences. We are looking for a senior, hands‑on Machine Learning Engineering Manager to lead a small team working at the intersection of model evaluation, agent optimization, and data generation. In this role, you will help define how evaluation closes the loop with model and product development, turning observed quality gaps into targeted improvements to prompts, agent harnesses, datasets, and models.You will combine technical depth with people leadership. You should be comfortable moving from research papers and experimental results to production‑quality ML pipelines, while mentoring engineers and aligning teams around a clear technical direction. Your work will span Apple Foundation Models and product teams, with the goal of creating repeatable evaluation and refinement loops that improve the quality of Apple intelligence experiences.

Description

As a Senior Machine Learning Engineering Manager in AIML Evaluation, you will lead the technical strategy and execution for agent evaluation and automatic optimization. You will own systems that evaluate foundation models and agents, diagnose failure modes, and use those signals to drive automated prompt, context, tool, rubric, and agent-harness improvements. You will also help establish the interfaces between evaluation and post‑training so that high‑value failures can be converted into targeted data, environments, reward signals, and measurable model improvements.This is a hands‑on leadership role. You will prototype new approaches, participate in architecture and code reviews, design experiments, and help your team translate emerging research into scalable evaluation and optimization pipelines. You will partner closely with Apple Foundation Models, product engineering teams, and other AIML groups to build an evaluation flywheel that connects real product behavior with model and agent refinement. You will also work across the organization to advance synthetic data generation for both evaluation and post‑training, with strong attention to data quality, representativeness, privacy, and reproducibility.

Responsibilities
  • Architects and builds scalable evaluation systems for foundation models and agents, including benchmarks, LLM‑based evaluators, simulation environments, trajectory analysis, and regression testing.
  • Establishes an end‑to‑end evaluation flywheel with Apple Foundation Models and product teams that connects observed failures to diagnosis, targeted refinement, post‑training, and measurable quality improvement.
  • Leads, mentors, and grows a small team of machine learning engineers while remaining deeply involved in technical design, experimentation, implementation, and review.
  • Defines the technical strategy and roadmap for automatic prompt, context, tool, rubric, and agent‑harness optimization for agentic development and model evaluation.
  • Develops methods that convert evaluation findings into actionable model‑improvement signals, including targeted datasets, synthetic trajectories, reward or preference signals, and optimization objectives.
  • Partners across AIML to design and scale synthetic data generation pipelines for evaluation and post‑training.
  • Applies and adapts recent research in LLM and agent evaluation, automatic optimization, LLM‑as‑judge, reward modeling, test‑time search, and post‑training to production‑quality workflows.
Minimum Qualifications
  • 8+ years of professional experience in machine learning, applied research, or software engineering, including experience building production ML systems or large‑scale experimentation platforms.
  • 3+ years of technical leadership experience, including direct people management of machine learning or software engineers and a demonstrated ability to mentor and grow strong technical talent.
  • Master’s or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field.
  • Strong hands‑on programming and software engineering skills, particularly in Python, with experience building reliable ML pipelines using modern machine learning or deep learning frameworks.
  • Deep experience with large language models or agentic systems, including evaluation of multi‑turn behavior, tool use, planning, reasoning, or other action‑taking workflows.
  • Experience building automated evaluation methods such as LLM‑based judges, rubrics, reward models, simulation‑based evaluation, or scalable benchmark infrastructure.
  • Experience with at least one model or agent refinement area such as automatic prompt or context optimization, post‑training, preference optimization, reinforcement learning, or agent‑harness optimization.
  • Excellent communication and collaboration skills, with demonstrated ability to align research, engineering, and product teams around ambiguous technical problems.
Preferred Qualifications
  • Track record of applying recent machine learning research to production systems or high‑impact product development.
  • Experience with automatic prompt or context optimization, agent‑search methods, evaluator optimization, or multi‑objective optimization for agentic systems.
  • Experience generating and evaluating synthetic datasets, tool‑use trajectories, or multi‑turn agent interactions, including methods for filtering, deduplication, diversity, and quality control.
  • Experience designing evaluation systems that combine offline benchmarks, simulation, human evaluation, and product‑or‑usage‑derived signals.
  • Experience with privacy‑preserving or on‑device machine learning and evaluation.
  • Demonstrated ability to influence technical strategy across organizational boundaries and communicate complex model‑quality tradeoffs to senior technical leaders.
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $237,600 and $401,700, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Apple Inc. • Seattle (WA), Northern (KY)

On-site
USD 226,000 - 382,000
AIML - Director, Data Science and Insights
AIML - Director, Data Science and Insights

Apple Inc. • Cupertino (CA)

On-site
USD 305,000 - 488,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase plan
+1
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
AIML - Sr Machine Learning Engineer, Data and ML Innovation
AIML - Sr Machine Learning Engineer, Data and ML Innovation

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Employee stock programs
Discretionary bonuses
Relocation assistance
+2
AIML - Data Scientist, Evaluation
AIML - Data Scientist, Evaluation

Apple • Cupertino (CA)

On-site
USD 147,000 - 273,000
Employee stock programs
Comprehensive medical coverage
Retirement benefits
+1
AIML - Machine Learning Researcher, Data and ML Innovation
AIML - Machine Learning Researcher, Data and ML Innovation

Apple Inc. • Santa Clara (CA)

On-site
USD 181,000 - 319,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services
+1
Algorithm Evaluation Manager
Algorithm Evaluation Manager

Apple Inc. • Sunnyvale (CA)

On-site
USD 206,000 - 356,000
Stock options
Discretionary bonuses
Relocation assistance
+1
AIML - Machine Learning Engineer, Foundation Models
AIML - Machine Learning Engineer, Foundation Models

Apple Inc. • Seattle (WA)

On-site
USD 184,700 - 324,800
Medical and dental coverage
Employee stock programs
Relocation assistance
LLM Machine Learning Engineer, Models and Agent Science, AIML
LLM Machine Learning Engineer, Models and Agent Science, AIML

Apple Inc. • Seattle (WA)

On-site
USD 185,000 - 325,000
Sr Engineering Program Manager, AIML Foundation Models
Sr Engineering Program Manager, AIML Foundation Models

Apple Inc. • Cupertino (CA)

On-site
USD 175,000 - 264,000