Multimodal ML Engineer - Vision, Audio & Text AI

AI Breaking Wire

Mountain View (CA)

Hybrid

USD 200,000 - 320,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Comprehensive health programs
Generous vacation & family leave
On-site gourmet cafeterias

Job summary

Google DeepMind is seeking a Machine Learning Engineer focused on multimodal models in Mountain View, CA. You will build and optimize models that fuse vision, audio, text, and video data, targeting low latency and high throughput at scale.

You should have 3+ years of industry experience with Python and deep learning frameworks (TensorFlow or PyTorch), plus a strong grasp of CV, NLP, and multimodal fusion techniques. Excellent collaboration skills are required.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, or a related field.
  • 3+ years of industry experience building, training, and deploying deep learning models at scale.
  • Strong proficiency in Python and deep learning frameworks such as TensorFlow or PyTorch.
  • Solid understanding of computer vision, natural language processing, and multimodal fusion techniques.
  • Excellent problem-solving and collaboration skills.

Responsibilities

  • Build and train cutting-edge multimodal AI models that integrate vision, audio, text, and video data.
  • Optimize model architectures for low latency and high throughput during large-scale inference.
  • Partner with product and research teams to translate foundational research into consumer-facing applications.
  • Conduct rigorous evaluations to measure model performance, fairness, and robustness across diverse domains.

Skills

python
pytorch
tensorflow
computer-vision
nlp
multimodal

Education

Bachelor's or Master's in CS/AI

Job description

Google DeepMind is seeking a Machine Learning Engineer focused on multimodal models in Mountain View, CA. You will build and optimize models that fuse vision, audio, text, and video data, targeting low latency and high throughput at scale.

You should have 3+ years of industry experience with Python and deep learning frameworks (TensorFlow or PyTorch), plus a strong grasp of CV, NLP, and multimodal fusion techniques. Excellent collaboration skills are required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Multimodal AI Engineer: Vision, Audio & Text
Multimodal AI Engineer: Vision, Audio & Text

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 210,000 - 310,000
Equity program
Healthcare
On-site gym
+3
Machine Learning Engineer, Multimodal Models
Machine Learning Engineer, Multimodal Models

AI Breaking Wire • Mountain View (CA)

Hybrid
USD 200,000 - 320,000
Comprehensive health programs
Generous vacation & family leave
On-site gourmet cafeterias
Machine Learning Engineer, Multimodal AI
Machine Learning Engineer, Multimodal AI

AI Breaking Wire • Mountain View (CA)

Hybrid
USD 220,000 - 360,000
Competitive compensation & equity
Health & wellness coverage
Hybrid work environment
+1
Applied AI Engineer, Multimodal
Applied AI Engineer, Multimodal

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 210,000 - 310,000
Equity program
Healthcare
On-site gym
+3
Multimodal AI Research Scientist
Multimodal AI Research Scientist

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Real-Time Multimodal AI/ML Engineer for Vision
Real-Time Multimodal AI/ML Engineer for Vision

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 150,000 - 278,000
Comprehensive medical and dental cover
Retirement benefits
Discounted products and free services
+3
Multimodal AI Research Engineer — Vision & LLMs
Multimodal AI Research Engineer — Vision & LLMs

Apple • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Applied Multimodal ML Research Engineer
Applied Multimodal ML Research Engineer

Apple Inc. • Sunnyvale (CA), Northern (KY)

Hybrid
USD 150,000 - 225,000
Generative AI Research Engineer — Multimodal, Prod-Ready
Generative AI Research Engineer — Multimodal, Prod-Ready

AI Breaking Wire • New York (NY), Northern (KY)

Hybrid
USD 220,000 - 340,000
Stock options
Medical, dental, vision
Learning & conference budget
Multimodal AI Engineer — Image/Video & Audio
Multimodal AI Engineer — Image/Video & Audio

xAI • Seattle (WA)

On-site
USD 180,000 - 440,000
Equity
Comprehensive medical, vision, and dental coverage
401(k) retirement plan
+2