Senior AI Evaluation & Benchmarking Lead

Apple

Cambridge

On-site

GBP 100,000 - 140,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Apple seeks a Senior Manager to lead evaluation development and data science for Siri. You will shape measurement systems across multiple Apple platforms, guiding teams, and collaborating with Finance, Annotation Ops, and international partners to ensure high-quality human-in-the-loop evaluations.

Join a cross-disciplinary team focused on robust metrics, privacy-first design, and scalable tooling to manage rich data assets while delivering state-of-the-art evaluation capabilities for users

Qualifications

  • Minimum Qualifications: Agentic coding proficiency to achieve data-science, data collection and visualization tasks.
  • Good understanding of metrics, crowd science, data collection, annotation analysis, statistics.
  • Ability to work independently and cross-functionally to integrate in partner team reporting systems and pipelines.
  • Excellent communication skills and the ability to thrive in a highly collaborative work environment.

Responsibilities

  • Design and execute efficient data collection processes using humans in the loop.
  • Lead annotation efforts for various languages and devices.
  • Plan and manage the budget of annotator resourcing with Finance, Annotation Ops, and International team partners.
  • Track and improve the quality of the human judgments through revision of annotator training materials, clear annotation questions, annotator training, and efficient reviewing mechanisms.
  • Collaborate with other engineering teams to design and build a tooling ecosystem for managing and browsing rich datasets.

Job description

Apple seeks a Senior Manager to lead evaluation development and data science for Siri. You will shape measurement systems across multiple Apple platforms, guiding teams, and collaborating with Finance, Annotation Ops, and international partners to ensure high-quality human-in-the-loop evaluations.

Join a cross-disciplinary team focused on robust metrics, privacy-first design, and scalable tooling to manage rich data assets while delivering state-of-the-art evaluation capabilities for users

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation & Benchmarking Lead
AI Evaluation & Benchmarking Lead

Apple Inc. • Cambridge

On-site
GBP 120,000 - 170,000
Benchmarking Project Lead, Siri Evaluation
Benchmarking Project Lead, Siri Evaluation

Apple • Cambridge

On-site
GBP 100,000 - 140,000
Senior Manager, Agentic AI Evaluation
Senior Manager, Agentic AI Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 180,000
Benchmarking Project Lead, Siri Evaluation
Benchmarking Project Lead, Siri Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 170,000
Sr Manager, Siri Agentic Evaluation
Sr Manager, Siri Agentic Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 180,000
Sr Manager, Siri Agentic Evaluation
Sr Manager, Siri Agentic Evaluation

APPLE • Cambridge

On-site
GBP 120,000 - 180,000
AI Evaluations Engineer
AI Evaluations Engineer

ConnexAI • Manchester

On-site
GBP 50,000 - 70,000
Part-Time AI Researcher for Benchmarking (MacBook Required)
Part-Time AI Researcher for Benchmarking (MacBook Required)

Mercor • Greater London

On-site
GBP 21,000 - 34,000
Senior AI Compute Systems Engineer – Benchmarking & Metrics
Senior AI Compute Systems Engineer – Benchmarking & Metrics

United States Digital Space LLC • West of England, Greater London

On-site
GBP 70,000 - 110,000
Flexible working
Generous leave
Pension matching
+2
Senior AI Safety Data Scientist & Evaluation Lead
Senior AI Safety Data Scientist & Evaluation Lead

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 120,000
Unlimited Annual Leave Policy
Private healthcare and dental
Enhanced parental leave
+3