Benchmarking Project Lead, Siri Evaluation

Apple

Cambridge

On-site

GBP 100,000 - 140,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Apple seeks a Senior Manager to lead evaluation development and data science for Siri. You will shape measurement systems across multiple Apple platforms, guiding teams, and collaborating with Finance, Annotation Ops, and international partners to ensure high-quality human-in-the-loop evaluations.

Join a cross-disciplinary team focused on robust metrics, privacy-first design, and scalable tooling to manage rich data assets while delivering state-of-the-art evaluation capabilities for users

Qualifications

  • Minimum Qualifications: Agentic coding proficiency to achieve data-science, data collection and visualization tasks.
  • Good understanding of metrics, crowd science, data collection, annotation analysis, statistics.
  • Ability to work independently and cross-functionally to integrate in partner team reporting systems and pipelines.
  • Excellent communication skills and the ability to thrive in a highly collaborative work environment.

Responsibilities

  • Design and execute efficient data collection processes using humans in the loop.
  • Lead annotation efforts for various languages and devices.
  • Plan and manage the budget of annotator resourcing with Finance, Annotation Ops, and International team partners.
  • Track and improve the quality of the human judgments through revision of annotator training materials, clear annotation questions, annotator training, and efficient reviewing mechanisms.
  • Collaborate with other engineering teams to design and build a tooling ecosystem for managing and browsing rich datasets.

Job description

Summary

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.Evaluation is at the heart of how we build our product. As Siri AI becomes more and more powerful and offers ever richer experiences to our users, our evaluations have to keep pace. We are seeking a senior manager to help drive our evaluation efforts. The role will involve managing teams working on evaluation development and data science, and leading high-impact initiatives to bring state-of-the-art agentic evaluation to the whole Siri team.

Description

As part of the work on next generation Siri, we are developing novel measurements of its quality. To ensure that the evaluation systems we are building are reliable, we plan benchmarking their accuracy on a wide range of features, locales, and platforms using humans in the loop.

Key Responsibilities

Design and execute efficient data collection processes using humans in the loopLead annotation efforts for various languages and devicesPlan and manage the budget of annotator resourcing with Finance, Annotation Ops, and International team partners.Track and improve the quality of the human judgements through revision of annotator training materials, clear annotation questions, annotator training, and efficient reviewing mechanismsCollaborate with other engineering teams to design and build a tooling ecosystem for managing and browsing rich datasets

Minimum Qualifications

Agentic Coding proficiency to achieve data-science, data collection and visualisation tasksGood understanding of metrics, crowd science, data collection, annotation analysis, statisticsAbility to work independently and cross-functionally to integrate in partner team reporting systems and pipelinesExcellent communication skills and the ability to thrive in a highly collaborative work environment

Preferred Qualifications

Attunement to computational linguistics, language quality, human in the loop evaluationGood engineering practices to create sustainable and easy to use data management pipelinesPython experience and other tools for data collection and visualisation

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Benchmarking Project Lead, Siri Evaluation
Benchmarking Project Lead, Siri Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 170,000
Senior AI Evaluation & Benchmarking Lead
Senior AI Evaluation & Benchmarking Lead

Apple • Cambridge

On-site
GBP 100,000 - 140,000
Sr Manager, Siri Agentic Evaluation
Sr Manager, Siri Agentic Evaluation

APPLE • Cambridge

On-site
GBP 120,000 - 180,000
Sr Manager, Siri Agentic Evaluation
Sr Manager, Siri Agentic Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 180,000
AI Evaluation & Benchmarking Lead
AI Evaluation & Benchmarking Lead

Apple Inc. • Cambridge

On-site
GBP 120,000 - 170,000
Senior Manager, Agentic AI Evaluation
Senior Manager, Agentic AI Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 180,000
AI Evaluations Engineer
AI Evaluations Engineer

ConnexAI • Manchester

On-site
GBP 50,000 - 70,000
Full Stack Software Engineer
Full Stack Software Engineer

Worky • Greater London

On-site
GBP 90,000 - 130,000
Software Engineer, iOS
Software Engineer, iOS

Meta • Greater London

Hybrid
GBP 85,000 - 135,000
Staff+ Software Engineer (RL Data Platform)
Staff+ Software Engineer (RL Data Platform)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 130,000
Health insurance
Equity options
Relocation support
+5