Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Socket.dev

Seattle (WA)

On-site

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking a Machine Learning Engineer focused on Evaluation & Insights for the Human-Centered AI team. You will bridge human perception with model performance, architect evaluation frameworks, and design MLOps pipelines to assess and improve large-scale AI systems.

You will work with data scientists, researchers, and engineers to ensure AI experiences are safe, reliable, and aligned with human expectations across features like search and recommendations.

Qualifications

  • 5+ years of relevant industry experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
  • Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Strong ability to interpret unstructured model outputs and synthesize qualitative findings into actionable guidance and training objectives.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Deep familiarity with AI quality metrics, hallucination detection techniques, and LLM-as-a-judge frameworks.
  • Experience building internal tools or automated pipelines for ML workflows using tools like MLflow, Weights & Biases, or similar platforms.
  • Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and Fine-Tuning.
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or related field

Responsibilities

  • Bridge the gap between human perception and algorithmic performance to evaluate AI systems.
  • Architect robust evaluation frameworks for foundation models and generative AI.
  • Design scalable MLOps pipelines for model assessment and monitoring.
  • Translate qualitative failure modes into programmatic guardrails and training signals (e.g., SFT, RLHF/DPO).
  • Collaborate with Software Engineering, Product, Research and Responsible AI teams to ensure reliable and safe AI experiences.

Skills

Python
PyTorch
JAX
Hugging Face
LLMs
NLP
Model evaluation
ML inference pipelines
RAG architectures
Fine-Tuning

Education

Bachelor's or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, or related technical field

Tools

MLflow
Weights & Biases
Vector databases
Semantic search

Job description

Imagine what you could do here. At Apple, great new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish! Are you passionate about music, movies, and the world of Artificial Intelligence and Machine Learning? So are we! Join our Human-Centered AI team for Apple Media Services. In this role, you'll represent the user perspective on new features, review and analyze data, and evaluate AI models powering everything from search and recommendations to other innovative features. You'll also collaborate with Data Scientists, Researchers, and Engineers to drive improvements across our platforms.

Description

We are looking for a Machine Learning Engineer focused on Evaluation & Insights for the Human-Centered AI team. In this role, you will bridge the gap between human perception and algorithmic performance, helping evaluate and optimize Foundation Models and generative AI systems. You will architect robust evaluation frameworks, design scalable MLOps pipelines for model assessment, and translate qualitative failure modes into programmatic guardrails and training signals (e.g., SFT, RLHF/DPO). This role blends deep ML engineering expertise with strong analytical judgment to assess, interpret, and improve the behavior of advanced AI models. You will work cross-functionally with Software Engineering, Product, Research and Responsible AI teams at Apple to ensure that our AI experiences are reliable, safe, and aligned with human expectations.

Minimum Qualifications
  • 5+ years of relevant industry experience in ML Engineering or Applied Research.
  • Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
  • Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Strong ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Deep familiarity with AI quality metrics, hallucination detection techniques (e.g., SelfCheckGPT), model alignment (RLHF/DPO), and LLM-as-a-judge frameworks (e.g., G-Eval, DeepEval).
  • Experience building internal tools or automated pipelines for ML workflows using tools like MLflow, Weights & Biases, or similar platforms.
  • Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and Fine-Tuning.
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field
Preferred Qualifications
  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Human-Centered AI ML Engineer — Evaluation & Insights
Human-Centered AI ML Engineer — Evaluation & Insights

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 190,000
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Medical and dental coverage
Employee stock programs
Stock purchase plan
+3
Human-Centered AI ML Engineer: Evaluations & Insights
Human-Centered AI ML Engineer: Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Sr. Applied Scientist, AI Evaluation & Quality Systems
Sr. Applied Scientist, AI Evaluation & Quality Systems

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Apple Benefits
Relocation assistance
Employee stock purchase plan
+2
ML Evaluation Specialist, Human Data
ML Evaluation Specialist, Human Data

Apple Inc. • Cupertino (CA)

On-site
USD 144,000 - 264,000
Comprehensive medical and dental coverage
Retirement benefits
Education reimbursement
AIML - Sr Engineering Specialist, Evaluation
AIML - Sr Engineering Specialist, Evaluation

Apple • Seattle (WA)

On-site
USD 130,000 - 180,000