Sr. Machine Learning Engineer, Speech LLM Evaluation

Socket.dev

Cupertino (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Apple is seeking a role focused on evaluating speech LLMs, building datasets that stress-test models, and defining metrics to turn outputs into reliable signals. You will work with modeling, infrastructure, and product teams to ensure rigorous evaluation before shipping to users across Apple platforms.

The position emphasizes real-time speech understanding, cross-functional collaboration, and a strong foundation in data-driven assessment within a privacy-conscious framework.

Qualifications

  • Bachelor's degree in a related field or equivalent practical experience.
  • Experience building or working with text, speech or audio evaluation pipelines and metrics.
  • Proficiency in Python and experience building data processing pipelines at scale.
  • Experience curating datasets for ML evaluation or training.
  • Working knowledge of statistics for model performance evaluation.
  • Familiarity with LLM evaluation techniques including automated and human evaluation methods.
  • Strong written and verbal communication skills.

Responsibilities

  • Own the data and metrics foundation for evaluating speech LLMs across accuracy, robustness, and conversational quality.
  • Build and curate evaluation datasets reflecting real usage and design metrics and automated judges.
  • Collaborate with modeling, infrastructure, and product partners to ensure quick, consistent, rigorous evaluation before customer release.

Skills

Python
Data pipelines
Statistics
LLM evaluation
Communication

Education

Bachelor's degree

Tools

Spark
Pandas

Job description

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and real-time conversational AI efforts, partnering directly with the modeling teams. We're growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether they're ready. You'll help define how Apple measures a new class of models that listen, speak, and reason. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.

Description

This role owns the data and metrics foundation for evaluating speech LLMs (e.g., real-time speech understanding and generation models) across accuracy, robustness, and conversational quality. You'll build and curate evaluation datasets that reflect real usage — from personalized named-entity queries to multi-turn fluid conversations — and design the metrics and automated judges that turn model outputs into actionable, trustworthy signal. You'll work closely with modeling, infrastructure, and product partners to make sure every new model is evaluated quickly, consistently, and at the right level of rigor before it reaches customers.

Minimum Qualifications
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.
  • Experience building or working with text, speech or audio evaluation pipelines and metrics.
  • Proficiency in Python and experience building data processing pipelines at scale.
  • Experience curating or annotating datasets for machine learning evaluation or training.
  • Working knowledge of statistics as applied to measuring model performance and interpreting evaluation results.
  • Familiarity with large language model evaluation techniques, including automated (LLM-as-judge) and human evaluation methods.
  • Strong written and verbal communication skills, with the ability to explain evaluation results to both technical and non-technical audiences.
Preferred Qualifications
  • Experience evaluating audio-native or multimodal (speech-in, speech-out) large language models.
  • Experience designing or running human evaluation studies (e.g., side-by-side comparisons, MOS ratings) at scale.
  • Familiarity with personalization and named-entity evaluation challenges in speech systems.
  • Experience with multilingual or international audio dataset development.
  • Experience with distributed data processing frameworks (e.g., Spark) for large-scale audio dataset generation.
  • Publication record or demonstrated contributions in speech, audio ML, or NLP evaluation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Speech LLM Evaluation Engineer — Real-Time Audio AI
Senior Speech LLM Evaluation Engineer — Real-Time Audio AI

Apple Inc. • Cupertino (CA)

Hybrid
USD 190,000 - 325,000
Medical and dental coverage
Apple stock programs
Tuition reimbursement
Sr. Machine Learning Engineer, Speech LLM Evaluation
Sr. Machine Learning Engineer, Speech LLM Evaluation

Apple Inc. • Cupertino (CA)

Hybrid
USD 190,000 - 325,000
Medical and dental coverage
Apple stock programs
Tuition reimbursement
Senior Speech AI Evaluation Engineer
Senior Speech AI Evaluation Engineer

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
Senior Applied Scientist, Multilingual AI Evaluation
Senior Applied Scientist, Multilingual AI Evaluation

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 230,000
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 240,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000
Sr. Machine Learning Engineer, Siri Speech
Sr. Machine Learning Engineer, Siri Speech

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000
Medical and dental coverage
Employee stock programs
Relocation assistance
AIML - Senior Machine Learning Research Engineer, LLM Post-training (Multilinguality)
AIML - Senior Machine Learning Research Engineer, LLM Post-training (Multilinguality)

Apple • New York (NY)

On-site
USD 180,000 - 240,000
Sr. Machine Learning Research Engineer, Siri Speech
Sr. Machine Learning Research Engineer, Siri Speech

Apple Inc. • Cupertino (CA)

On-site
USD 185,000 - 325,000