Research Scientist, Gemini Audio i18n, DeepMind

Google

Mountain View (CA)

On-site

USD 174,000 - 252,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Google DeepMind in Mountain View, CA seeks a Research Scientist for Gemini Audio i18n to advance multilingual audio models. You will design large-scale experiments, prototype architectures, and evaluate A2A speech systems across languages.

You will publish findings and contribute to open collaboration, with strong Python/C++ coding, DL frameworks, and a track record in speech or LLM research.

Qualifications

  • Bachelor's degree or equivalent in CS, Speech Recognition, or related field.
  • Experience in research or development of speech-related models.
  • Proficient in Python or C++ and DL frameworks.
  • Experience with audio data and multilingual datasets.

Responsibilities

  • Create large-scale evaluation sets for multilingual A2A models.
  • Prototype novel architectures to improve audio understanding and generation.
  • Collaborate with cross-functional teams to deploy solutions.

Skills

Python
C++
Speech Recognition
LLMs
Multilingual datasets

Education

Bachelor's degree in CS or related

Tools

PyTorch
JAX
TensorFlow

Job description

Research Scientist, Gemini Audio i18n, DeepMind

Mountain View, CA, USA; New York, NY, USA

Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Mountain View, CA, USA; New York, NY, USA.

Minimum qualifications:
  • Bachelor's degree in Computer Science, Speech Recognition, Computational Linguistics, a related technical field, or equivalent practical experience.
  • Experience conducting research or development in Speech Recognition, Text-to-Speech (TTS), or Large Language Models (LLMs).
  • Experience coding in Python or C++ and using deep learning frameworks such as PyTorch, JAX, or TensorFlow.
  • Experience working with audio data, speech processing, or multilingual datasets.
Preferred qualifications:
  • Master's degree or Ph.D. in Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Computational Linguistics, or a related field.
  • 3 years of experience in Large Language Models (LLMs) or multimodal foundation models.
  • Experience developing audio-to-audio (A2A) architectures, end-to-end speech models, or spoken dialog systems.
  • Experience scaling speech models across international languages, accents, or low-resource locales.
  • Publication record in speech or machine learning conferences (e.g., ICASSP, INTERSPEECH, NeurIPS, or ACL).
About the job

As an organization, Google maintains a portfolio of research projects driven by fundamental research, new product innovation, product contribution and infrastructure goals, while providing individuals and teams the freedom to emphasize specific types of work. As a Research Scientist, you'll setup large-scale tests and deploy promising ideas quickly and broadly, managing deadlines and deliverables while applying the latest theories to develop new and improved products, processes, or technologies. From creating experiments and prototyping implementations to designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more.

As a Research Scientist, you'll also actively contribute to the wider research community by sharing and publishing your findings, with ideas inspired by internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits. Learn more about benefits at Google.

Responsibilities
  • Create comprehensive evaluation sets and benchmarks to measure audio-to-audio (A2A) model performance across international languages, accents, and regional dialects.
  • Propose, prototype, and evaluate novel modeling techniques to improve A2A audio understanding, dialog, and audio generation capabilities with a focus on scalability.
  • Identify performance gaps in current multilingual audio models and collaborate with cross-functional research and engineering teams to deploy solutions.

Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

If you have a need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, Gemini Audio i18n, DeepMind
Research Scientist, Gemini Audio i18n, DeepMind

Google Inc. • New York (NY)

On-site
USD 174,000 - 252,000
Research Scientist, Gemini Audio i18n, DeepMind
Research Scientist, Gemini Audio i18n, DeepMind

Google • New York (NY)

Hybrid
USD 174,000 - 252,000
Research Scientist, Gemini Audio i18n, DeepMind
Research Scientist, Gemini Audio i18n, DeepMind

Google DeepMind • New York (NY)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Comprehensive benefits
Research Scientist, Gemini Audio i18n, DeepMind
Research Scientist, Gemini Audio i18n, DeepMind

Google DeepMind • Mountain View (CA)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Benefits
Audio to Audio Research Scientist, DeepMind
Audio to Audio Research Scientist, DeepMind

Google • United States

On-site
USD 207,000 - 300,000
Audio to Audio Research Scientist, DeepMind
Audio to Audio Research Scientist, DeepMind

Google • New York (NY)

Hybrid
USD 207,000 - 300,000
Equity
Benefits
Research Engineer, Winslow, DeepMind
Research Engineer, Winslow, DeepMind

Google • Mountain View (CA)

On-site
USD 174,000 - 253,000
ML Research Scientist, Audio Algorithms
ML Research Scientist, Audio Algorithms

Google Inc. • Mountain View (CA), Irvine (CA)

On-site
USD 174,000 - 252,000
Equity
Benefits
Research Engineer, Conversational Agentic AI, DeepMind
Research Engineer, Conversational Agentic AI, DeepMind

Google Inc. • New York (NY)

On-site
USD 207,000 - 300,000
Senior Research Scientist, Beam
Senior Research Scientist, Beam

Google • Mountain View (CA)

On-site
USD 174,000 - 252,000