Machine Learning Engineer, Speech/Audio

attention labs

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

attention labs is building the audio modeling stack in a fast-moving startup setting in San Francisco. You will own end-to-end lifecycle from data to production inference, shipping models that recognize who is being spoken to in real time.

You will turn research into dependable inference under real-world noise and latency constraints, design evaluation harnesses, and collaborate with founders to shape next steps for the product's audio intelligence.

Qualifications

  • Strong applied ML experience with audio, speech, or signal-heavy models.
  • Ability to own a model from data collection to production inference, not just notebooks.
  • Fluency with modern deep learning tooling such as PyTorch and rigorous measurement.

Responsibilities

  • Own the audio model lifecycle end to end: data, features, architecture, training, evaluation, and production inference.
  • Build and improve models behind addressee detection and real-time decisions on who is being spoken to.
  • Turn research prototypes into fast, dependable inference under real-world noise and accents.
  • Design evaluation harnesses and datasets to honestly measure impact of changes.
  • Work directly with founders and early customers to shape next model objectives.

Skills

Audio/speech ML
Production model lifecycle
PyTorch
Latency awareness

Tools

PyTorch

Job description

San Francisco · Full-time · Machine Learning

You will own the audio modeling stack end to end, from raw signal to production inference. You are the person who ships the models that decide, in real time, who a voice agent is being spoken to.

  • Own the audio model lifecycle end to end: data, features, architecture, training, evaluation, and the inference path that runs in production.
  • Build and improve the models behind addressee detection, deciding whether speech is meant for the agent or for someone else in the room.
  • Turn research prototypes into fast, dependable inference that holds up under real world noise, overlap, and accents.
  • Design the evaluation harness and datasets that tell us, honestly, whether a change made the product better.
  • Work directly with founders and early customers to shape what the model needs to do next.
What we are looking for
  • Strong applied machine learning experience with audio, speech, or other signal-heavy models.
  • Comfort owning a model from data collection through to something running in production, not just a notebook.
  • Fluency with modern deep learning tooling such as PyTorch, and the discipline to measure before and after every change.
  • A bias toward shipping, and toward the smallest experiment that answers the question.
  • Enough systems sense to care about latency, memory, and how a model behaves under load.
  • Bonus: experience with real time audio, source separation, or on-device inference.
About attention labs

attention labs is early: a small team defining a new category at the intersection of speech, cognitive neuroscience, and machine learning. We work from San Francisco, Toronto, and Memphis, and we are remote-friendly for the right person.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Audio Modeling Engineer — Real-Time Production Inference
Audio Modeling Engineer — Real-Time Production Inference

attention labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Embedded Systems Engineer, On-Device Inference
Embedded Systems Engineer, On-Device Inference

attention labs • Memphis (TN), Northern (KY)

Hybrid
USD 110,000 - 150,000
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research Scientist, Addressee Detection / Speech Separation
Research Scientist, Addressee Detection / Speech Separation

attention labs • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Machine Learning Intern
Machine Learning Intern

Bland AI • San Francisco (CA)

On-site
USD 40,000 - 65,000
Competitive intern compensation
Mentorship from researchers on front‑f
All the tools you need
+2
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
ML Engineer
ML Engineer

Catalyst Labs • New York (NY)

On-site
USD 120,000 - 140,000
Competitive compensation
Bonus opportunities
Equity in the company
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
Member of Technical Staff - Multi-Modal, Audio San Francisco · Boston · Hybrid
Member of Technical Staff - Multi-Modal, Audio San Francisco · Boston · Hybrid

Liquid AI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
100% health premiums
401(k) matching up to 4%
Unlimited PTO
+1
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific Global Solutions, Inc. • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Competitive compensation
Hybrid/Remote options
GPU infrastructure access
+1