Human-Centered AI Evaluation Engineer

Apple Inc.

Cupertino (CA)

On-site

USD 184,700 - 324,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple Inc. is seeking an Evaluation & Insights Machine Learning Engineer in Cupertino, CA to lead comprehensive model evaluations for LLMs and multimodal models, building evaluation frameworks and translating findings into actionable improvements for product and engineering teams.

You will collaborate with Data Scientists, Researchers, and Engineers to ensure AI systems are reliable, safe, and aligned with human expectations, while advancing prompt engineering, RAG strategies, and ML CI/CD

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field

Responsibilities

  • Lead Rigorous Model Evaluations: Architect and execute comprehensive evaluation suites for LLMs and multimodal models, identifying edge cases in multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
  • Advanced Scoring Frameworks: Develop deterministic, heuristic, and LLM-assisted evaluation frameworks (e.g., LLM-as-a-judge, reward modeling) to quantify human-perceived quality metrics (e.g., helpfulness, hallucination rates).
  • Actionable Signal Extraction: Translate qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for model training and inference.
  • Improve Performance: Partner with engineering teams to refine model behavior, leveraging evaluation telemetry to inform prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning.
  • Latent Pattern Recognition: Apply advanced ML techniques (e.g., embedding-based clustering, representation learning, perturbation analysis) to systematically map error taxonomies and latent failure manifolds in model outputs.
  • MLOps & Automation: Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines.
  • Distributed Evaluation Pipelines: Architect scalable, distributed inference and processing pipelines (e.g., Ray, vLLM) for high-throughput model evaluation, automated annotation, and output analysis at scale.
  • Human-Centric Metrics: Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability.
  • Auto-Evaluator Systems: Build automated evaluation pipelines utilizing LLMs to assess outputs at scale, optimizing for high correlation with human baseline annotations.
  • Cross-Functional Partnership: Collaborate with ML researchers, software developers, and product managers across Apple to translate product requirements into scalable, reliable, and efficient model evaluation infrastructure.

Skills

Python
PyTorch
JAX
Hugging Face
LLM evaluation
Prompt engineering
RAG architectures
Model evaluation
NLP systems
Data analysis

Education

Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field

Tools

MLflow
Weights & Biases

Job description

Apple Inc. is seeking an Evaluation & Insights Machine Learning Engineer in Cupertino, CA to lead comprehensive model evaluations for LLMs and multimodal models, building evaluation frameworks and translating findings into actionable improvements for product and engineering teams.

You will collaborate with Data Scientists, Researchers, and Engineers to ensure AI systems are reliable, safe, and aligned with human expectations, while advancing prompt engineering, RAG strategies, and ML CI/CD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer: AI Evaluation & LLM Systems
ML Engineer: AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Senior AI Safety & Evaluation Engineer
Senior AI Safety & Evaluation Engineer

Apple Inc. • Seattle (WA)

On-site
USD 184,700 - 324,800
Employee stock purchase plan
Relocation assistance
Education reimbursement
Responsible AI ML Engineer: Safety & Evaluation
Responsible AI ML Engineer: Safety & Evaluation

Apple • California (MO)

On-site
USD 180,000 - 230,000
Senior AI/ML Evaluation Engineer
Senior AI/ML Evaluation Engineer

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 115,000 - 236,000
Stock programs
Restricted Stock Units (RSU)
Medical and dental coverage
+4
Senior AIML Engineer — AI Model Evaluation & Benchmarking
Senior AIML Engineer — AI Model Evaluation & Benchmarking

Apple Inc. • Cupertino (CA)

On-site
USD 212,000 - 387,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer - AI & ML Evaluation Frameworks
Machine Learning Engineer - AI & ML Evaluation Frameworks

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services
AIML - AI Software Engineer, Evaluation
AIML - AI Software Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
ML Engineer — Health AI Evaluation & Safety Frameworks
ML Engineer — Health AI Evaluation & Safety Frameworks

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Discounted products and free services