Sr. Machine Learning Engineer, Speech LLM Evaluation

Apple Inc.

Cambridge (MA)

Hybrid

USD 166,000 - 292,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
Medical and dental coverage
Tuition reimbursement
Relocation assistance
Employee stock purchase plan

Job summary

Apple Inc. Cambridge is seeking a Sr. Machine Learning Engineer to lead evaluation of audio LLMs in the Speech domain.

You will design datasets, craft metrics, and build scalable pipelines to stress-test personalized and multilingual ASR/TTs, integrating with modeling and product teams. The role emphasizes rigorous evaluation, collaboration with infrastructure, and translating results for stakeholders. Join a team advancing how Apple measures and improves listening and conversational AI across

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • Experience building or working with text, speech or audio evaluation pipelines and metrics.
  • Proficiency in Python and building data processing pipelines at scale.
  • Experience curating or annotating datasets for ML evaluation or training.
  • Working knowledge of statistics for evaluating model performance.
  • Familiarity with LLM evaluation techniques including automated and human evaluation.
  • Strong written and verbal communication to explain results to technical and non-technical audiences.

Responsibilities

  • Designs and curates audio evaluation datasets representing real-world usage, including personalized, multilingual, and conversational scenarios.
  • Defines and implements evaluation metrics for audio LLMs across accuracy, robustness, and conversational/generation quality.
  • Builds automated evaluation pipelines and LLM-as-judge tooling to scale audio model assessment.
  • Analyzes evaluation results to identify gaps, regressions, and opportunities for improvement.
  • Partners with human-evaluation programs to design rating protocols and validate automated metrics against human judgment.
  • Collaborates with infrastructure teams to integrate new evaluation sets into shared tooling.
  • Contributes evaluation methodology for new audio LLM capabilities as models evolve.

Skills

Python
Evaluation pipelines
Dataset curation
Statistics
LLM evaluation methods
Communication
Spark

Education

Bachelor's degree

Tools

Spark

Job description

Sr. Machine Learning Engineer, Speech LLM Evaluation

Cambridge, Massachusetts, United States Machine Learning and AI

Join the team redefining what a deeply personal and integrated assistant can be.As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and real-time conversational AI efforts, partnering directly with the modeling teams. We're growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether they're ready. You'll help define how Apple measures a new class of models that listen, speak, and reason.This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.

Description

This role owns the data and metrics foundation for evaluating speech LLMs (e.g., real-time speech understanding and generation models) across accuracy, robustness, and conversational quality. You'll build and curate evaluation datasets that reflect real usage — from personalized named-entity queries to multi-turn fluid conversations — and design the metrics and automated judges that turn model outputs into actionable, trustworthy signal. You'll work closely with modeling, infrastructure, and product partners to make sure every new model is evaluated quickly, consistently, and at the right level of rigor before it reaches customers.

Responsibilities
  • Designs and curates audio evaluation datasets that represent real-world usage, including personalized, multilingual, and conversational scenarios.
  • Defines and implements evaluation metrics for audio LLMs, spanning accuracy, robustness, and conversational/generation quality.
  • Builds automated evaluation pipelines and LLM-as-judge tooling to scale audio model assessment without sacrificing reliability.
  • Analyzes model evaluation results to identify accuracy gaps, regressions, and opportunities for hillclimbing, and communicates findings to modeling teams.
  • Partners with human-evaluation programs to design rating protocols and validate that automated metrics correlate with human judgment.
  • Collaborates with infrastructure teams to integrate new evaluation sets and metrics into shared tooling.
  • Contributes evaluation methodology for new audio LLM capabilities as they emerge, adapting existing frameworks to novel model behaviors.
Minimum Qualifications
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.
  • Experience building or working with text, speech or audio evaluation pipelines and metrics.
  • Proficiency in Python and experience building data processing pipelines at scale.
  • Experience curating or annotating datasets for machine learning evaluation or training.
  • Working knowledge of statistics as applied to measuring model performance and interpreting evaluation results.
  • Familiarity with large language model evaluation techniques, including automated (LLM-as-judge) and human evaluation methods.
  • Strong written and verbal communication skills, with the ability to explain evaluation results to both technical and non-technical audiences.
Preferred Qualifications
  • Experience evaluating audio-native or multimodal (speech-in, speech-out) large language models.
  • Experience designing or running human evaluation studies (e.g., side-by-side comparisons, MOS ratings) at scale.
  • Familiarity with personalization and named-entity evaluation challenges in speech systems.
  • Experience with multilingual or international audio dataset development.
  • Experience with distributed data processing frameworks (e.g., Spark) for large-scale audio dataset generation.
  • Publication record or demonstrated contributions in speech, audio ML, or NLP evaluation.

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $165,800 and $292,400, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation.

  • Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
  • Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
  • Comprehensive medical and dental coverage,
  • Retirement benefits,
  • A range of discounted products and free services,
  • For formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition.
  • Discretionary bonuses or commission payments
  • Relocation.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Machine Learning Engineer, Speech LLM Evaluation
Sr. Machine Learning Engineer, Speech LLM Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 190,000 - 325,000
Medical and dental coverage
Apple stock programs
Tuition reimbursement
Sr. Machine Learning Engineer, Speech LLM Evaluation
Sr. Machine Learning Engineer, Speech LLM Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Sr. Machine Learning Engineer, Siri Speech
Sr. Machine Learning Engineer, Siri Speech

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Medical and dental coverage
Employee stock programs
Relocation assistance
Machine Learning Architect - Conversational Speech
Machine Learning Architect - Conversational Speech

Apple Inc. • Cupertino (CA)

On-site
USD 263,000 - 394,000
Comprehensive medical and dental cover
Retirement benefits
Discounted products and free services
+2
AIML - Machine Learning Researcher, Speech
AIML - Machine Learning Researcher, Speech

Apple Inc. • Cupertino (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Medical and dental coverage
Retirement benefits
Apple stock programs
+2
Machine Learning Manager, Siri Speech
Machine Learning Manager, Siri Speech

Apple Inc. • Cupertino (CA)

On-site
USD 206,000 - 310,000
Discretionary bonuses
Relocation
Stock programs
+2
Sr. Machine Learning Scientist, Siri Speech
Sr. Machine Learning Scientist, Siri Speech

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Stock programs
RSU awards
Tuition reimbursement
+1
Machine Learning Engineer, Search & Knowledge Quality
Machine Learning Engineer, Search & Knowledge Quality

Apple Inc. • Seattle (WA)

On-site
USD 150,400 - 277,600
Stock options
Medical and dental coverage
Relocation assistance
+1
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Machine Learning Engineer, Human Centered AI - Evaluations & Insights

Apple Inc. • Seattle (WA)

On-site
USD 142,000 - 263,000
Machine Learning Engineer, Search & Knowledge Quality
Machine Learning Engineer, Search & Knowledge Quality

Apple Inc. • Santa Clara (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock purchase program
+2