Machine Learning Engineer - Agentic AI Evaluation Frameworks

Apple Inc.

Cupertino, Northern (CA, KY)

Hybrid

USD 185,000 - 278,000

Full time

27 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical and dental coverage
Retirement benefits
Employee stock purchase plan
Relocation assistance

Job summary

Apple Inc. in Cupertino, CA is seeking a Machine Learning Engineer to build evaluation frameworks for LLM-powered experiences across the Commerce domain.

You will design automated evaluation pipelines, create golden datasets, and calibrate evaluators with human feedback to drive product quality.

Collaborate with ML, software, product, data science, and UX teams to ensure scalable, reliable AI that meets user needs worldwide.

Qualifications

  • Experience in ML engineering, ML evaluation, or related fields.
  • Strong Python programming and production software development experience.
  • Experience with LLMs, Generative AI, or NLP.
  • Experience designing automated evaluation frameworks, metrics, benchmarks.
  • Understanding of LLM architectures, prompting, and retrieval-augmented generation.
  • Experience with Human-in-the-Loop evaluation and data-quality workflows.
  • Strong statistical analysis and experimental design.
  • Bachelor's in CS/ML/AI or related field, or equivalent experience.

Responsibilities

  • Design and build scalable evaluation systems for LLM, Generative AI, and agentic products.
  • Develop evaluation datasets, benchmarks, and quality signals.
  • Define metrics across accuracy, relevance, groundedness, and task completion.
  • Build reusable evaluation infrastructure, dashboards, and tooling.
  • Integrate evaluation into CI/CD workflows and release readiness.

Skills

Python
ML Engineering
Evaluation frameworks
NLP/LLM
Software engineering
Data pipelines
Statistical analysis

Education

Bachelor's degree in CS/ML/AI
Master's degree preferred

Tools

CI/CD pipelines

Job description

Machine Learning Engineer - Agentic AI Evaluation Frameworks

Cupertino, California, United States Machine Learning and AI

Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish.The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.In this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle.You will help define how we measure the quality of AI experiences across the Commerce domain, including Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI, and Platform Self-Service.This is an opportunity to work at the intersection of machine learning, software engineering, data, and product quality, helping ensure our AI experiences are accurate, relevant, grounded, reliable, and useful for users around the world.

Description

As a Machine Learning Evaluation Engineer, you will design and build scalable evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI products.

  • Design and develop automated evaluation frameworks and pipelines for AI-powered products.
  • Define evaluation methodologies and quality metrics across dimensions such as accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion.
  • Build and maintain high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets.
  • Develop Auto Eval capabilities that enable teams to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences.
  • Design and implement model-based evaluation approaches, including LLM-as-a-Judge, while developing appropriate calibration and validation methodologies.
  • Develop Human-in-the-Loop (HITL) evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient.
  • Define evaluation rubrics, annotation guidelines, grading criteria, and quality standards in partnership with product teams, domain experts, and annotation teams.
  • Build mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and reliability.
  • Evaluate end-to-end AI systems, including retrieval, context construction, prompts, model responses, tool use, APIs, and downstream product experiences.
  • Develop evaluation methodologies for multi-turn conversations, personalization, recommendations, tool use, reasoning, and agentic task execution.
  • Perform detailed error analysis and failure-mode investigation to identify opportunities for model, prompt, retrieval, dataset, and product improvements.
  • Build reusable evaluation infrastructure, APIs, dashboards, and developer tooling that can scale across multiple AI products and teams.
  • Integrate evaluation into development and CI/CD workflows, enabling automated regression detection, quality gates, and release-readiness assessments.
  • Connect offline evaluation results with production signals to continuously improve evaluation coverage and product quality.
  • Partner closely with Machine Learning, Software Engineering, Product, Quality Engineering, Human Interface, and Data Science teams throughout research, development, evaluation, launch, and continuous improvement.
Minimum Qualifications
  • Typically requires a minimum of 7 years of related experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.
  • Strong programming skills in Python and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.
  • Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.
  • Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
  • Understanding of modern LLM application architectures, including prompting, embeddings, retrieval-augmented generation (RAG), tool use, and agentic workflows.
  • Experience with model-based evaluation techniques and an understanding of the strengths and limitations of approaches such as LLM-as-a-Judge.
  • Experience with Human-in-the-Loop evaluation, annotation, or data-quality workflows.
  • Strong understanding of statistical analysis, experimentation, sampling, and measurement methodologies.
  • Experience performing model error analysis, failure analysis, and root-cause investigation.
  • Ability to work effectively across Machine Learning, Engineering, Product, Quality, and Data teams.
  • Excellent written and verbal communication skills, with the ability to translate complex technical findings into clear, actionable recommendations.
  • Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, Data Science, Statistics, Electrical Engineering, or a related technical field, or equivalent industry experience.
Preferred Qualifications
  • Experience building evaluation infrastructure for production-scale LLM or Generative AI applications.
  • Experience evaluating RAG, conversational systems, AI agents, personalization, recommendations, or multimodal AI.
  • Experience building golden datasets, regression suites, automated quality gates, and continuous evaluation pipelines.
  • Experience integrating ML evaluation into CI/CD and production release processes.
  • Experience with prompt evaluation, model comparison, experiment tracking, and AI observability.
  • Experience evaluating multilingual AI experiences across languages, locales, and markets.
  • Familiarity with responsible AI evaluation, including robustness, safety, bias, and adversarial testing.
  • Experience developing internal ML platforms, developer tooling, or self-service evaluation capabilities used across multiple teams.
  • Experience working with large-scale datasets and distributed ML or data-processing infrastructure.
  • Master's degree in Computer Science, Machine Learning, Artificial Intelligence, Data Science, Statistics, Electrical Engineering, or a related technical field, or equivalent industry experience.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $277,600, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Sr Machine Learning Engineering Manager, Evaluation
AIML - Sr Machine Learning Engineering Manager, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 238,000 - 402,000
Stock options
Relocation assistance
Comprehensive medical/dental coverage
+1
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Stock purchase plan
+3
Sr. Applied Scientist, AI Evaluation & Quality Systems
Sr. Applied Scientist, AI Evaluation & Quality Systems

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Apple Benefits
Relocation assistance
Employee stock purchase plan
+2
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Sr Engineering Program Manager, Evaluation - Special Projects
Sr Engineering Program Manager, Evaluation - Special Projects

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 176,000 - 312,000
Discretionary bonuses/stock options
Employee stock purchase plan
Relocation assistance
+2
Senior AI Engineer - Services Special Projects
Senior AI Engineer - Services Special Projects

Apple Inc. • San Francisco (CA)

On-site
USD 185,000 - 325,000
Medical and dental coverage
Retirement benefits
Employee stock purchase plan
+4
AIML - Software Engineer - AI, Evaluation
AIML - Software Engineer - AI, Evaluation

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 278,000
Comprehensive medical and dental cover
Retirement benefits
Employee stock purchase plan
+1
Senior Engineering Manager, Agentic AI & Intelligent Reasoning
Senior Engineering Manager, Agentic AI & Intelligent Reasoning

Apple Inc. • Seattle (WA)

On-site
USD 268,000 - 402,000
Senior ML Platform Engineer - Agentic Systems - Special Projects
Senior ML Platform Engineer - Agentic Systems - Special Projects

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 185,000 - 325,000
Employee stock programs
Relocation support
Education reimbursement
LLM Machine Learning Engineer, Models and Agent Science, AIML
LLM Machine Learning Engineer, Models and Agent Science, AIML

Apple Inc. • Seattle (WA)

On-site
USD 185,000 - 325,000