Speech & Audio ML Research Engineer

Hume AI

New York (NY)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hume AI is seeking a Research Engineer to develop speech-language and audio models, train at scale, and rigorously evaluate methods. You will work with the research team to turn questions into experiments, train and refine models, and build software that supports distributed training, inference, and benchmarking.

You will collaborate with customer teams to align research objectives with real-world data and ensure reproducible, high-quality results across in-house and client projects.

Qualifications

  • At least two years of experience contributing to model training or fine-tuning with large-scale datasets (text, audio, image, or video).
  • Strong Python and PyTorch skills, including developing, debugging, and maintaining training and evaluation code.
  • Experience writing robust research systems that go beyond notebooks or prototypes.
  • Strong experimental judgment: baselines, evaluations, and careful interpretation of results.
  • Evidence of research ability via publications, model releases, or substantive open-source work.
  • Comfort iterating quickly on uncertain directions and owning open-ended technical problems.

Responsibilities

  • Investigate research questions in speech-language, audio, and multimodal ML; design experiments to test hypotheses.
  • Train, fine-tune, validate, and further develop models in-house and for customer projects.
  • Build reliable software for distributed model training, inference, and benchmarking; ensure reproducibility.
  • Develop data workflows for large-scale datasets used in training and evaluation.
  • Design evaluation protocols and benchmarks to measure quality, generalization, and robustness.
  • Analyze datasets, model outputs, and failures to guide next steps in development.
  • Collaborate with customer teams to define questions, understand data, and conduct training/evaluation.
  • Document experimental designs, results, and technical decisions for reproducibility.

Skills

Python
PyTorch
Model training
Experimentation
Multimodal ML
Distributed systems
Research publications

Tools

Git
Jupyter
Linux

Job description

Hume AI is seeking a Research Engineer to develop speech-language and audio models, train at scale, and rigorously evaluate methods. You will work with the research team to turn questions into experiments, train and refine models, and build software that supports distributed training, inference, and benchmarking.

You will collaborate with customer teams to align research objectives with real-world data and ensure reproducible, high-quality results across in-house and client projects.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Hume AI • New York (NY)

On-site
USD 120,000 - 180,000
Senior AI Research Engineer, Speech & Audio Models
Senior AI Research Engineer, Speech & Audio Models

Humeai • New York (NY)

On-site
USD 180,000 - 350,000
Senior/Staff AI Research Engineer
Senior/Staff AI Research Engineer

Humeai • New York (NY)

On-site
USD 180,000 - 350,000
Research Scientist - Audio
Research Scientist - Audio

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Senior Audio ML Research Engineer
Senior Audio ML Research Engineer

David AI • San Francisco (CA)

On-site
USD 210,000 - 360,000
Unlimited PTO
Health, dental, vision coverage
FSA & HSA access
+3
Remote Research Engineer — AI Audio & Models
Remote Research Engineer — AI Audio & Models

ElevenLabs • Town of Poland (NY)

On-site
USD 140,000 - 190,000
Senior Software Engineer - Microservices
Senior Software Engineer - Microservices

Hume AI • New York (NY)

On-site
USD 180,000 - 240,000
Applied ML Engineer: From Research to Production
Applied ML Engineer: From Research to Production

deepgram • United States

On-site
USD 140,000 - 180,000
Senior Speech ML Engineer — End-to-End Voice Systems
Senior Speech ML Engineer — End-to-End Voice Systems

Cantina • United States

Remote
USD 200,000 - 220,000
Equity
Healthcare
PTO 42 days
+5
Speech ML Engineer - Large-Scale Audio Models | Hybrid + Equity
Speech ML Engineer - Large-Scale Audio Models | Hybrid + Equity

Plaud • San Francisco (CA)

Hybrid
USD 195,000 - 365,000
Equity
Health coverage
401(k) matching
+3