Machine Learning Engineer, Multimodal Models

AI Breaking Wire

Mountain View (CA)

Hybrid

USD 200,000 - 320,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Comprehensive health programs
Generous vacation & family leave
On-site gourmet cafeterias

Job summary

Google DeepMind is seeking a Machine Learning Engineer focused on multimodal models in Mountain View, CA. You will build and optimize models that fuse vision, audio, text, and video data, targeting low latency and high throughput at scale.

You should have 3+ years of industry experience with Python and deep learning frameworks (TensorFlow or PyTorch), plus a strong grasp of CV, NLP, and multimodal fusion techniques. Excellent collaboration skills are required.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, or a related field.
  • 3+ years of industry experience building, training, and deploying deep learning models at scale.
  • Strong proficiency in Python and deep learning frameworks such as TensorFlow or PyTorch.
  • Solid understanding of computer vision, natural language processing, and multimodal fusion techniques.
  • Excellent problem-solving and collaboration skills.

Responsibilities

  • Build and train cutting-edge multimodal AI models that integrate vision, audio, text, and video data.
  • Optimize model architectures for low latency and high throughput during large-scale inference.
  • Partner with product and research teams to translate foundational research into consumer-facing applications.
  • Conduct rigorous evaluations to measure model performance, fairness, and robustness across diverse domains.

Skills

python
pytorch
tensorflow
computer-vision
nlp
multimodal

Education

Bachelor's or Master's in CS/AI

Job description

# Machine Learning Engineer, Multimodal ModelsGoogle DeepMind## Job Description### About Google DeepMindGoogle DeepMind brings together brilliant minds from all over the world to solve hard problems in computer science and advance the state of the art in artificial intelligence. We are hiring a Machine Learning Engineer to focus on multimodal models.### Responsibilities- Build and train cutting-edge multimodal AI models that seamlessly integrate vision, audio, text, and video data.- Optimize model architectures for low latency and high throughput during large-scale inference.- Partner with product and research teams to transition foundational research into consumer-facing applications.- Conduct rigorous evaluations to measure model performance, fairness, and robustness across diverse domains.### Requirements- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, or a related field.- 3+ years of industry experience building, training, and deploying deep learning models at scale.- Strong proficiency in Python and deep learning frameworks such as TensorFlow or PyTorch.- Solid understanding of computer vision, natural language processing, and multimodal fusion techniques.- Excellent problem-solving and collaboration skills.### Benefits- Competitive salary, bonus, and equity packages.- Comprehensive health and wellness programs.- Generous vacation policy and paid family leave.- On-site gourmet cafeterias and wellness facilities.## Skills & Tagspythonpytorchtensorflowcomputer-visionnlpmultimodal## Job DetailsFull-timeMountain View, CA$200k – $320k USDPosted September 21, 2026Expires November 20, 2026
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Multimodal AI
Machine Learning Engineer, Multimodal AI

AI Breaking Wire • Mountain View (CA)

Hybrid
USD 220,000 - 360,000
Competitive compensation & equity
Health & wellness coverage
Hybrid work environment
+1
Applied AI Engineer, Multimodal
Applied AI Engineer, Multimodal

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 210,000 - 310,000
Equity program
Healthcare
On-site gym
+3
Multimodal ML Engineer - Vision, Audio & Text AI
Multimodal ML Engineer - Vision, Audio & Text AI

AI Breaking Wire • Mountain View (CA)

Hybrid
USD 200,000 - 320,000
Comprehensive health programs
Generous vacation & family leave
On-site gourmet cafeterias
Multimodal AI Engineer: Vision, Audio & Text
Multimodal AI Engineer: Vision, Audio & Text

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 210,000 - 310,000
Equity program
Healthcare
On-site gym
+3
Research Engineer, Generative AI
Research Engineer, Generative AI

AI Breaking Wire • New York (NY), Northern (KY)

Hybrid
USD 220,000 - 340,000
Stock options
Medical, dental, vision
Learning & conference budget
Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

Pantera Capital • San Francisco (CA)

On-site
USD 150,000 - 230,000
Machine Learning: Multimodal Foundation Models
Machine Learning: Multimodal Foundation Models

The Bot Company • San Francisco (CA)

On-site
USD 180,000 - 290,000
Machine Learning Engineer - Multimodal Intelligence
Machine Learning Engineer - Multimodal Intelligence

Socket.dev • Seattle (WA)

On-site
USD 150,000 - 210,000
AI Research Engineer
AI Research Engineer

AI Breaking Wire • Northern (KY), New York (NY)

Hybrid
USD 210,000 - 310,000
Competitive salary
Health insurance
Dental and vision insurance
+2
Research, Vision Expertise
Research, Vision Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3