Benchmarking Project Lead, Siri Evaluation

Apple Inc.

Cambridge

On-site

GBP 120,000 - 170,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Apple Inc. in Cambridge, England is seeking a Senior Manager to lead evaluation development and data science for Siri AI, shaping state-of-the-art agentic evaluation across iOS, iPadOS, macOS, watchOS and visionOS.

You will manage teams, drive high-impact initiatives, and collaborate with Finance and International partners to plan budgets and resourcing. This role emphasizes privacy-first design and real-world impact across platforms.

Qualifications

  • Agentic coding proficiency for data science, data collection and visualization.
  • Strong understanding of metrics, crowd science, data collection, annotation analysis, statistics.
  • Ability to work independently and cross-functionally with partner team reporting systems.
  • Excellent communication skills and the ability to thrive in a highly collaborative work environment.

Responsibilities

  • Design and execute data collection processes using humans in the loop.
  • Lead annotation efforts for various languages and devices.
  • Plan and manage annotator resourcing budgets with Finance, Annotation Ops, and International team partners.
  • Track and improve quality of human judgements through training materials and review processes.
  • Collaborate with engineering teams to build tooling for managing and browsing rich datasets.

Skills

Agentic coding
Metrics understanding
Crowd science
Data collection
Annotation analysis
Statistics
Independent work
Cross-functional collaboration
Communication skills

Tools

Python
Data visualization tools

Job description

Selection changes the language of the page/content

Cambridge, England, United Kingdom Machine Learning and AI

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.Evaluation is at the heart of how we build our product. As Siri AI becomes more and more powerful and offers ever richer experiences to our users, our evaluations have to keep pace. We are seeking a senior manager to help drive our evaluation efforts. The role will involve managing teams working on evaluation development and data science, and leading high-impact initiatives to bring state-of-the-art agentic evaluation to the whole Siri team.

Description

As part of the work on next generation Siri, we are developing novel measurements of its quality. To ensure that the evaluation systems we are building are reliable, we plan benchmarking their accuracy on a wide range of features, locales, and platforms using humans in the loop.

Responsibilities
  • Design and execute efficient data collection processes using humans in the loop
  • Lead annotation efforts for various languages and devices
  • Plan and manage the budget of annotator resourcing with Finance, Annotation Ops, and International team partners.
  • Track and improve the quality of the human judgements through revision of annotator training materials, clear annotation questions, annotator training, and efficient reviewing mechanisms
  • Collaborate with other engineering teams to design and build a tooling ecosystem for managing and browsing rich datasets
Minimum Qualifications
  • Agentic Coding proficiency to achieve data-science, data collection and visualisation tasks
  • Good understanding of metrics, crowd science, data collection, annotation analysis, statistics
  • Ability to work independently and cross-functionally to integrate in partner team reporting systems and pipelines
  • Excellent communication skills and the ability to thrive in a highly collaborative work environment
Preferred Qualifications
  • Attunement to computational linguistics, language quality, human in the loop evaluation
  • Good engineering practices to create sustainable and easy to use data management pipelines
  • Python experience and other tools for data collection and visualisation

At Apple, we're not all the same. And that's our greatest strength. We draw on the differences in who we are, what we've experienced and how we think. Because to create products that serve everyone, we believe in including everyone. Therefore, we are committed to treating all applicants fairly and equally. As a registered Disability Confident employer, we will work with applicants to make any reasonable accommodations. Apple will consider for employment all qualified applicants with criminal backgrounds in a manner consistent with applicable law. Learn more

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Benchmarking Project Lead, Siri Evaluation
Benchmarking Project Lead, Siri Evaluation

Apple • Cambridge

On-site
GBP 100,000 - 140,000
Sr Manager, Siri Agentic Evaluation
Sr Manager, Siri Agentic Evaluation

Apple Inc. • Cambridge

On-site
GBP 120,000 - 180,000
Sr Manager, Siri Agentic Evaluation
Sr Manager, Siri Agentic Evaluation

APPLE • Cambridge

On-site
GBP 120,000 - 180,000
Senior AI Evaluation & Benchmarking Lead
Senior AI Evaluation & Benchmarking Lead

Apple • Cambridge

On-site
GBP 100,000 - 140,000
AI Evaluation & Benchmarking Lead
AI Evaluation & Benchmarking Lead

Apple Inc. • Cambridge

On-site
GBP 120,000 - 170,000
Principal Machine Learning Engineer, AI & Data Platforms (AiDP)
Principal Machine Learning Engineer, AI & Data Platforms (AiDP)

Apple • Greater London

On-site
GBP 95,000 - 130,000
Engineering Manager, ML Infrastructure, London
Engineering Manager, ML Infrastructure, London

Apple Inc. • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Data Science Manager – Strategic Data Solutions
Senior Data Science Manager – Strategic Data Solutions

Apple Inc. • Cambridge

On-site
GBP 142,000 - 214,000
Comprehensive medical and dental cover
Retirement benefits
Discounted products and free services
+2
ML Software Engineer, London
ML Software Engineer, London

Apple Inc. • Greater London

Hybrid
GBP 120,000 - 180,000
Developer Relations, Technology Evangelist
Developer Relations, Technology Evangelist

Apple Inc. • Greater London

On-site
GBP 85,000 - 110,000