Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc.

Seattle (WA)

On-site

USD 142,000 - 263,000

Full time

17 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. is seeking a Machine Learning Engineer for the Human Centered AI team in Seattle to lead evaluation and insights for AI systems powering Apple Media Services.

The role focuses on bridging human perception and model performance, architecting evaluation frameworks, and translating failure modes into training signals. Collaboration spans ML Researchers, Software Engineers, and Product teams.

Qualifications

  • 5+ years of experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
  • Proven experience building scalable ML inference pipelines and model evaluation workflows.

Responsibilities

  • Lead rigorous model evaluations for LLMs and multimodal models.
  • Develop evaluation frameworks to quantify human-perceived quality metrics.
  • Translate qualitative failure signals into actionable data and training signals.

Skills

Python
PyTorch
JAX
Hugging Face

Education

Bachelor’s or Master’s degree in Computer Science / ML / AI / Cognitive Science

Tools

MLflow
Weights & Biases

Job description

Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Seattle, Washington, United States Machine Learning and AI

Imagine what you could do here. At Apple, great new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish! Are you passionate about music, movies, and the world of Artificial Intelligence and Machine Learning? So are we!Join our Human-Centered AI team for Apple Media Services. In this role, you'll represent the user perspective on new features, review and analyze data, and evaluate AI models powering everything from search and recommendations to other innovative features. You'll also collaborate with Data Scientists, Researchers, and Engineers to drive improvements across our platforms.

Description

We are looking for a Machine Learning Engineer focused on Evaluation & Insights for the Human-Centered AI team. In this role, you will bridge the gap between human perception and algorithmic performance, helping evaluate and optimize Foundation Models and generative AI systems. You will architect robust evaluation frameworks, design scalable MLOps pipelines for model assessment, and translate qualitative failure modes into programmatic guardrails and training signals (e.g., SFT, RLHF/DPO).This role blends deep ML engineering expertise with strong analytical judgment to assess, interpret, and improve the behavior of advanced AI models. You will work cross-functionally with Software Engineering, Product, Research and Responsible AI teams at Apple to ensure that our AI experiences are reliable, safe, and aligned with human expectations.

Responsibilities
  • Lead Rigorous Model Evaluations: Architect and execute comprehensive evaluation suites for LLMs and multimodal models, identifying edge cases in multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
  • Advanced Scoring Frameworks: Develop deterministic, heuristic, and LLM-assisted evaluation frameworks (e.g., LLM-as-a-judge, reward modeling) to quantify human-perceived quality metrics (e.g., helpfulness, hallucination rates).
  • Actionable Signal Extraction: Translate qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for model training and inference.
  • Improve Performance: Partner with engineering teams to refine model behavior, leveraging evaluation telemetry to inform prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning.
  • Latent Pattern Recognition: Apply advanced ML techniques (e.g., embedding-based clustering, representation learning, perturbation analysis) to systematically map error taxonomies and latent failure manifolds in model outputs.
  • MLOps & Automation: Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines.
  • Distributed Evaluation Pipelines: Architect scalable, distributed inference and processing pipelines (e.g., Ray, vLLM) for high-throughput model evaluation, automated annotation, and output analysis at scale.
  • Human-Centric Metrics: Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability.
  • Auto-Evaluator Systems: Build automated evaluation pipelines utilizing LLMs to assess outputs at scale, optimizing for high correlation with human baseline annotations.
  • Cross-Functional Partnership: Collaborate with ML researchers, software developers, and product managers across Apple to translate product requirements into scalable, reliable, and efficient model evaluation infrastructure.
Minimum Qualifications
  • 5+ years of relevant industry experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
  • Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Strong ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Deep familiarity with AI quality metrics, hallucination detection techniques (e.g., SelfCheckGPT), model alignment (RLHF/DPO), and LLM-as-a-judge frameworks (e.g., G-Eval, DeepEval).
  • Experience building internal tools or automated pipelines for ML workflows using tools like MLflow, Weights & Biases, or similar platforms.
  • Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and Fine-Tuning.
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field
Preferred Qualifications
  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $142,300 and $263,300, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here - in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AIML - Sr Machine Learning Engineering Manager, Evaluation
AIML - Sr Machine Learning Engineering Manager, Evaluation

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 238,000 - 402,000
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Stock purchase plan
+3
Sr. Applied Scientist, AI Evaluation & Quality Systems
Sr. Applied Scientist, AI Evaluation & Quality Systems

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Apple Benefits
Relocation assistance
Employee stock purchase plan
+2
AIML - Machine Learning Engineer - Computer Vision & Audio, MIND
AIML - Machine Learning Engineer - Computer Vision & Audio, MIND

Apple Inc. • Seattle (WA)

On-site
USD 142,300 - 263,300
ML Evaluation Specialist, Human Data
ML Evaluation Specialist, Human Data

Apple Inc. • Cupertino (CA)

On-site
USD 144,000 - 264,000
Comprehensive medical and dental coverage
Retirement benefits
Education reimbursement
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)

Apple Inc. • Seattle (WA)

On-site
USD 175,000 - 263,300
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Apple Inc. • Seattle (WA), Northern (KY)

On-site
USD 226,000 - 382,000
Algorithm Evaluation Manager
Algorithm Evaluation Manager

Apple Inc. • Sunnyvale (CA)

On-site
USD 206,000 - 356,000
Stock options
Discretionary bonuses
Relocation assistance
+1
AIML - Sr Machine Learning Engineer, Data and ML Innovation
AIML - Sr Machine Learning Engineer, Data and ML Innovation

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 278,000
Employee stock programs
Discretionary bonuses
Relocation assistance
+2