Research Engineer

Kalpa Labs

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kalpa Labs is hiring a Founding ML Research Engineer to drive end-to-end development of generalist audio models. You will work across data, pre-training, post-training and evaluation, delivering research ideas to production rapidly with a small, compute-heavy team.

You’ll explore multi-modal architectures, scalable data pipelines, and high-performance ML systems. This role demands deep knowledge of large neural networks and strong software engineering skills to ship work quickly.

Qualifications

  • Experience researching pre-training and post-training large neural networks.
  • Strong ML systems and engineering depth (distributed training, performance, reliability).
  • Ability to spec, build, debug, and ship in ambiguity.

Responsibilities

  • Research and design efficient multi-modal architectures and codecs for audio.
  • Post-train audio models to enable instruction following and in-context learning for text and audio.
  • Build large-scale speech model pre-training and post-training pipelines (SFT/RLHF-style, distillation, etc.).
  • Create scalable data compute pipelines: dataset curation, filtering, tokenization, evaluation harnesses.

Skills

ML research
Distributed training
Multimodal models
Speech/audio
Hardware acceleration

Education

PhD in ML / AI

Tools

PyTorch
TensorFlow
HuggingFace

Job description

We’re hiring a Founding ML Research Engineer to work work through the full stack across data, pre-training, post-training & evals for training Generalist Audio Models. You’ll work through the entire stack with small team, tons of compute, high autonom, and see your research ideas making it to production within a week(s).

What you’ll do
  • Research better multi-modal architectures & codecs that are efficient across both spoken speech & general audio.
  • Post-train audio models to have LLM like instruction following & in-context learning but over both text and audio.
  • Build large-scale speech model pre-training and post-training (SFT/RLHF-style, distillation, preference optimization, etc.).
  • Build scalable data compute pipelines: dataset curation, filtering, mixing, tokenization/feature pipelines, evaluation harnesses.
  • Look at lots of data & hear lots of audio.
What we’re looking for
  • Industry/Academia experience pre-training / post-training large neural networks; speech/audio is a plus but not required, language/vision experience is also relevant.
  • Strong ML systems and engineering depth (distributed training, performance, reliability).
  • Comfort operating in ambiguity: you can spec, build, debug, and ship.
  • A hunger to always ask – what would the next frontier look like?
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - Audio [33361]
Research Scientist - Audio [33361]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Machine Learning Research Engineer
Senior Machine Learning Research Engineer

Carnaby Fox • San Francisco (CA)

On-site
USD 250,000 - 400,000
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Audio Expertise
Research, Audio Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Founding ML Research Engineer, Audio & Multimodal Models
Founding ML Research Engineer, Audio & Multimodal Models

Kalpa Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Audio ML Research Engineer
Senior Audio ML Research Engineer

David AI • San Francisco (CA)

On-site
USD 210,000 - 360,000
Unlimited PTO
Health, dental, vision coverage
FSA & HSA access
+3
ML Researcher, Speech
ML Researcher, Speech

DeepRec.ai • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Healthcare
Dental
+1
Research Scientist, Multilingual Audio & A2A Models
Research Scientist, Multilingual Audio & A2A Models

Google DeepMind • New York (NY)

On-site
USD 174,000 - 252,000
Equity
Bonus target
Comprehensive benefits
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000