Founding Machine Learning Engineer

Socket.dev

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Socket.dev in San Francisco is seeking a lead ML scientist to own the central problem of turning noisy physiological signals into continuous language. You will shape data collection, signal representation, model families, and evaluations, working closely with founders and sensing/ hardware leads.

We seek PhD-level research ability and production-quality coding speed, with the goal of building a product-ready, low-latency system that generalizes across people, sessions, and devices.

Qualifications

  • Experience in deep learning for speech and time-series data.
  • Strong research background with ability to publish or present results.
  • Ability to write production-grade, scalable code.

Responsibilities

  • Own the end-to-end ML problem from data to deployment.
  • Define data collection, signal representation, model families, and evaluation.
  • Scale training across thousands of hours and users.
  • Ensure real-time latency with continuous running systems.

Skills

Deep learning for speech
Time-series analysis
Biosignals
Neuroscience
BCIs
Research ability
Production-grade code

Education

PhD in a relevant field

Tools

Python
PyTorch
CUDA
Distributed GPU training

Job description

We are looking for the person who will own the central machine learning problem at Subvocal: turning weak, noisy, highly variable physiological signals into continuous language.

We have already built prototypes that decode subvocal speech at more than 200 words per minute. The much harder problem now is generalization. A model that works on one person, in one session, with one device placement is not a product. It needs to work when the same person returns the next day, when the hardware moves slightly, and eventually when a completely new person puts it on and gives us only a few minutes of calibration data.

You will lead that effort end to end. You will work directly with the founders and our sensing and hardware leads to decide what data we collect, how we represent the signal, which model families we pursue, and how we evaluate whether we are actually making progress.

Some of the problems you will work on include:

  • Learning general representations from large amounts of unlabeled and weakly labeled physiological time-series data.
  • Building continuous sequence models using approaches such as Conformers, Mamba-style architectures, CTC, transducers, and pretrained speech or language models.
  • Separating speech-related information from anatomy, placement, session, device, and environmental variation.
  • Adapting a large cross-user model to a new person from roughly 15 minutes of calibration data.
  • Designing honest user-held-out, session-held-out, and device-held-out evaluations.
  • Scaling training across thousands of hours and thousands of people.
  • Getting the complete system to run continuously with low enough latency for real-time use.

Our current ML stack is primarily Python, PyTorch, CUDA, and distributed GPU training, with custom infrastructure for signal processing, data collection, experiment tracking, and evaluation.

You might be a great fit if you have unusually strong experience in deep learning for speech (ASR), time series, biosignals, neuroscience, BCIs, radar, or another domain where signals are noisy and data distributions shift constantly. We are especially interested in people with PhD-level research ability, whether or not that came through a formal PhD, who are also comfortable writing production-quality code and moving quickly when the research direction changes.

This is not a role where you will be handed a model architecture and asked to improve it incrementally. You will help decide how the problem should be framed in the first place, build the initial ML organization around you, and directly determine whether this technology becomes a real product.

This is a full-time, in-person role in San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding ML Engineer: Generalizable Subvocal Speech
Founding ML Engineer: Generalizable Subvocal Speech

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 250,000
Founding Sensing Research Engineer
Founding Sensing Research Engineer

Socket.dev • San Francisco (CA)

On-site
USD 120,000 - 190,000
Applied AI Scientist, Language + ContextSan Francisco
Applied AI Scientist, Language + ContextSan Francisco

Stealth Neurotechnology Company • San Francisco (CA)

On-site
USD 180,000 - 230,000
Stock options
Comprehensive benefits package
401(k) program with matching
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Founding Machine Learning Engineer
Founding Machine Learning Engineer

Orbit • San Francisco (CA)

On-site
USD 225,000 - 275,000
Sr. Machine Learning Engineer
Sr. Machine Learning Engineer

Twenty80 LLC • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Engineer
ML Engineer

Catalyst Labs • New York (NY)

On-site
USD 120,000 - 140,000
Competitive compensation
Bonus opportunities
Equity in the company
Machine Learning Intern
Machine Learning Intern

Bland AI • San Francisco (CA)

On-site
USD 40,000 - 65,000
Competitive intern compensation
Mentorship from researchers on front‑f
All the tools you need
+2
ML Engineer - Data
ML Engineer - Data

Nuance Labs • Seattle (WA)

On-site
USD 90,000 - 120,000
Applied AI Scientist, Language + Context
Applied AI Scientist, Language + Context

Echo Neurotechnologies • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation, including stock options
Comprehensive benefits package
401(k) program with matching contributions