Senior Applied Scientist, Multilingual AI Evaluation

Socket.dev

Seattle (WA)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple seeks a highly skilled researcher to ensure AI features perform across languages and cultures. You will drive multilingual evaluation development from design to implementation, building scalable methods with measurement scientists and ML researchers.

You will publish novel findings while shipping practical evaluation tooling in Python across languages and scripts. You will shape evaluation methodology, test across dialects and cultures, and contribute to scalable pipelines that enable

Qualifications

  • Advanced linguistics knowledge with fluency in multiple languages beyond English.
  • Strong programming in Python for research tooling.
  • Experience evaluating/shipping features across languages or locales.
  • Ability to design rigorous benchmarks and evaluation protocols.

Skills

Linguistics
Python
Multilingual evaluation
Benchmark design
Statistical rigor
Cross-functional collaboration

Education

MS in Linguistics/Computational Linguistics/NLP

Tools

PyTorch
JAX

Job description

DESCRIPTION

In this role, you’ll help ensure Apple’s AI features work well across languages and cultures. Your goal is to make our evaluation tooling multilingual from the start so that engineers building AI features can design, test, and ship across the world from day one. It’s a broad applied science role: you’ll shape how Apple evaluates AI wherever the hardest questions are, and you’ll have the opportunity to publish novel work. The scientific challenge is real. How do we ensure we consistently evaluate AI features across different grammar, script, or cultural norms and how do we do this at scale? You’ll bring linguistic judgment to questions like these and, working with measurement scientists and ML researchers, turn it into validated methodology that holds across dozens of languages. This is a hands-on role. You’ll design and implement your own methods in Python, working closely with research and engineering partners, while staying focused on the science of getting evaluation right.


MINIMUM QUALIFICATIONS


  • MS in Linguistics, Computational Linguistics, NLP, Computer Science, or a related field — or equivalent research/work experience. Deep expertise in linguistics, with working fluency in the structure of multiple languages beyond English. Strong proficiency in Python. Solid understanding of LLMs and AI evaluation fundamentals, including how language models process and generate across languages. Demonstrated experience shipping or evaluating features across multiple languages or locales. Experience designing benchmarks, datasets, or human evaluation protocols, with attention to statistical rigor and reproducibility. Ability to drive initiatives independently and collaborate across a cross-functional, interdisciplinary team. Strong written and verbal communication skills.


PREFERRED QUALIFICATIONS


  • PhD in Linguistics, Computational Linguistics, or NLP with a focus on multilingual or cross-lingual modeling. Publications in NLP, multilingual evaluation, or evaluation methodology. Hands-on experience with modern ML frameworks (PyTorch, JAX) and with fine-tuning or evaluating LLMs. Experience with low-resource languages, dialectal variation, or sociolinguistics. Familiarity with localization/internationalization workflows and quality assessment. Experience with LLM-as-judge approaches, rubric design, or bias and fairness evaluation across languages. Fluency or professional proficiency in one or more languages in addition to English.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Multilingual AI Evaluation Scientist
Senior Multilingual AI Evaluation Scientist

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer - AI Evaluation & LLM Systems
Machine Learning Engineer - AI Evaluation & LLM Systems

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 230,000
AIML - Sr Engineering Specialist, Evaluation
AIML - Sr Engineering Specialist, Evaluation

Apple • Seattle (WA)

On-site
USD 130,000 - 180,000
AIML - Sr Machine Learning Engineer, Evaluation
AIML - Sr Machine Learning Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 212,000 - 387,000
Medical and dental coverage
Retirement benefits
Employee stock programs
+2
Senior AI Engineer - Services Special Projects
Senior AI Engineer - Services Special Projects

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 270,000
AIML - Sr Manager, Evaluation - Data Science & Insights
AIML - Sr Manager, Evaluation - Data Science & Insights

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 226,000 - 382,000
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation
AIML - Sr Applied AI Scientist - GenAI Model Autograding, Evaluation

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
AIML - AI Software Engineer, Evaluation
AIML - AI Software Engineer, Evaluation

Apple Inc. • Cupertino (CA)

On-site
USD 147,000 - 273,000
Comprehensive medical and dental coverage
Retirement benefits
Employee stock programs
+1
AIML - Sr Engineering Specialist, Evaluation
AIML - Sr Engineering Specialist, Evaluation

Apple Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 115,000 - 236,000
Stock programs
Restricted Stock Units (RSU)
Medical and dental coverage
+4
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 184,000 - 325,000