ML Evaluation Scientist: Foundation & Multimodal Models

Apple Inc.

Sunnyvale (CA)

On-site

USD 150,400 - 277,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical and dental cover
Retirement benefits
Relocation assistance

Job summary

Apple Inc. is seeking an expert in evaluating machine learning and deep learning models, including foundation models and multimodal systems.

You will craft robust evaluation frameworks, using traditional statistics and LLMs as judges to assess tasks like summarization and multimodal generation. The role requires strong Python expertise and a deep understanding of statistical methods, data quality, and model robustness, collaborating with ML engineers, data scientists, and ML infrastructure teams

Qualifications

  • BS and a minimum of 3 years relevant industry experience.
  • Strong experience evaluating supervised, unsupervised, and deep learning models.
  • Hands-on experience evaluating LLMs and using them as scoring/judging mechanisms.
  • Familiarity with multimodal models and related evaluation challenges.
  • Proficiency in Python and libraries such as NumPy, pandas, scikit-learn, PyTorch, or TensorFlow.
  • Solid understanding of statistical testing, sampling, confidence intervals, and metrics (e.g., precision/recall, BLEU, ROUGE, FID).
  • Strong documentation skills, including the ability to write technical reports and present to non-technical audiences.

Responsibilities

  • Develop robust methodologies to assess foundation models across diverse tasks.
  • Leverage LLMs as judges for subjective evaluations and multimodal tasks.
  • Build, curate, and lead evaluation datasets and benchmarks.
  • Collaborate with research, engineering, and product teams to align goals with user experience.
  • Conduct failure analysis to improve model robustness and document findings.

Skills

Python
NumPy
pandas
scikit-learn
PyTorch
TensorFlow
statistical testing
confidence intervals
BLEU/ROUGE/FID
documentation

Education

BS in a relevant field

Tools

OpenEval
ELO-based ranking
LLM-as-a-Judge frameworks

Job description

Apple Inc. is seeking an expert in evaluating machine learning and deep learning models, including foundation models and multimodal systems.

You will craft robust evaluation frameworks, using traditional statistics and LLMs as judges to assess tasks like summarization and multimodal generation. The role requires strong Python expertise and a deep understanding of statistical methods, data quality, and model robustness, collaborating with ML engineers, data scientists, and ML infrastructure teams

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Research Engineer - Multimodal Data & Foundation Models
ML Research Engineer - Multimodal Data & Foundation Models

shefsolutionsllc • Cupertino (CA)

Hybrid
USD 143,000 - 264,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+1
ML Engineer: AI Evaluation & LLM Systems
ML Engineer: AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
Foundation Models Evaluation Scientist (Multimodal ML)
Foundation Models Evaluation Scientist (Multimodal ML)

Apple • Boulder (CO)

On-site
USD 132,000 - 245,000
Machine Learning - Data Scientist
Machine Learning - Data Scientist

Apple • Boulder (CO)

On-site
USD 132,000 - 245,000
Comprehensive medical and dental coverage
Retirement benefits
Educational reimbursement
+1
Machine Learning - Data Scientist
Machine Learning - Data Scientist

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,400 - 277,600
Comprehensive medical and dental cover
Retirement benefits
Relocation assistance
Human-Centered AI Evaluation Engineer
Human-Centered AI Evaluation Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000
Multimodal LLMs Research Engineer
Multimodal LLMs Research Engineer

Apple Inc. • Sunnyvale (CA)

On-site
USD 150,000 - 278,000
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
LLM ML Engineer: Models & Agent Science
LLM ML Engineer: Models & Agent Science

Apple Inc. • Cupertino (CA)

On-site
USD 180,000 - 250,000
Health & dental
Retirement plan
Education reimbursement
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000