Audio AI Researcher: Multimodal, Low-Latency Modeling

Mosaic.tech

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training, post-training, and product to develop models that understand and generate audio with high fidelity and low latency.

The role requires deep experimentation, code writing, and collaboration with researchers, engineers, and designers to push the foundations of how AI learns and communicates.

Qualifications

  • Ability to design, run, and analyze experiments with empirical rigor.
  • Understanding of machine learning fundamentals and distributed compute environments.
  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Clarity in communication, an ability to explain complex technical concepts in writing.

Responsibilities

  • Own research projects on audio training, low‑latency inference and conversational responsiveness.
  • Design and train large-scale models that natively support audio input and output.
  • Investigate scaling behavior such as how data, model size, and compute affect capability and efficiency.
  • Build and maintain audio data pipelines, including preprocessing, filtering, segmentation, and alignment for training and evaluation.
  • Collaborate with data and infrastructure teams to scale audio training efficiently across distributed systems.
  • Publish and present research that moves the entire community forward. Share code, datasets, and insights that accelerate progress across industry and academia.

Skills

Experiment design
ML fundamentals
Python & DL frameworks
Communication skills

Education

Bachelor’s degree or equivalent
PhD (preferred) in CS/ML/Physics/Math

Tools

PyTorch
TensorFlow
JAX

Job description

Thinking Machines Lab in San Francisco is actively researching audio capabilities, blending rigorous theory with practical engineering to build multimodal AI systems. You will work across pre-training, post-training, and product to develop models that understand and generate audio with high fidelity and low latency.

The role requires deep experimentation, code writing, and collaboration with researchers, engineers, and designers to push the foundations of how AI learns and communicates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Audio AI Research for Multimodal Systems
Lead Audio AI Research for Multimodal Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Multimodal AI Research Scientist
Multimodal AI Research Scientist

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Senior Audio AI Researcher: Multimodal & Foundation Models
Senior Audio AI Researcher: Multimodal & Foundation Models

Via Licensing Corporation • Atlanta (GA)

On-site
USD 137,000 - 169,000
Audio-to-Audio AI Research Scientist
Audio-to-Audio AI Research Scientist

Google • United States

On-site
USD 207,000 - 300,000
Senior Audio-to-Audio AI Research Scientist
Senior Audio-to-Audio AI Research Scientist

Google Inc. • Mountain View (CA), New York (NY)

On-site
USD 207,000 - 300,000
Senior Applied ML Researcher: Video & Audio AI
Senior Applied ML Researcher: Video & Audio AI

Apple Inc. • Cupertino (CA)

On-site
USD 184,700 - 324,800
Multimodal AI Researcher: Generative Models, Realtime Vision
Multimodal AI Researcher: Generative Models, Realtime Vision

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000
Multimodal ML Researcher — Frontiers in AI
Multimodal ML Researcher — Frontiers in AI

Apple Inc. • Cambridge (MA)

On-site
USD 166,000 - 250,000