Machine Learning Intern

Uncover

San Francisco, Northern (CA, KY)

Hybrid

USD 34,000 - 55,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Mentorship from researchers
Office in Levi's Plaza, SF
Return offer potential

Job summary

Bland in San Francisco invites a Machine Learning Research Intern to own a focused project across our voice stack—from speech-to-text to text-to-speech. You’ll work with our research team on problems they are pursuing, not a busywork assignment, aiming for results worth shipping or publishing.

You will train and evaluate models on real telephony data, use distributed GPU infrastructure, and collaborate with engineers to move promising results toward production.

Qualifications

  • Currently pursuing a MS or PhD in ML, CS, EE, or related field, or equivalent research experience.
  • Comfortable reading a paper and reimplementing it without hand-holding.
  • Experience with self-supervised, generative, or multimodal modeling.
  • Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.
  • Strong intuition for audio quality and what makes synthetic speech sound wrong.

Responsibilities

  • Own a research question end to end.
  • Take one well-scoped problem from literature review through implementation, experimentation, and results.
  • Design ablations that isolate what actually caused an improvement.
  • Present your findings to the research team and defend the methodology.
  • Work on real systems.
  • Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
  • Use our distributed GPU infrastructure rather than toy-scale setups.
  • Where the result warrants it, work with engineers to move it toward production.

Skills

PyTorch
Experimentation
Speech audio models
Self-supervised modeling
Audio quality intuition

Education

MS or PhD in ML/CS/EE

Tools

GPU clusters

Job description

The Role: Machine Learning Research Intern, Audio

As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy. We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.

What You Will Do
  • Own a research question end to end
  • Take one well-scoped problem from literature review through implementation, experimentation, and results.
  • Design ablations that isolate what actually caused an improvement.
  • Present your findings to the research team and defend the methodology.
  • Work on real systems
  • Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
  • Use our distributed GPU infrastructure rather than toy-scale setups.
  • Where the result warrants it, work with engineers to move it toward production.
Choose your depth

Depending on your background and interests, your project may focus on:

  • Expressive and controllable text-to-speech, including prosody and emotion modeling
  • Neural audio codecs and discrete or continuous speech representations
  • ASR robustness for telephony, accents, and code switching
  • Real-time and streaming inference under latency constraints
  • Full-duplex conversation and turn-taking dynamics
What Makes You a Great Fit
Research foundations
  • Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.
  • Comfortable reading a paper and reimplementing it without hand-holding.
  • Experience with self-supervised, generative, or multimodal modeling.
  • Audio or speech grounding
  • Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.
  • Strong intuition for audio quality and what makes synthetic speech sound wrong.
  • Prior publications or open source contributions in speech or language AI are a strong signal, though not required.
Engineering ability
  • Fluent in PyTorch and comfortable in a real codebase.
  • Able to run your own experiments on GPU clusters without waiting to be unblocked.
How You Show Up
  • You identify the single experiment that validates an idea in days, not months.
  • You measure everything and let data drive decisions.
  • You are honest about negative results, because they are how we narrow the search.
  • You are obsessed with making voice agents sound truly human.
  • You use AI tools aggressively to amplify your own impact.
Benefits
  • Competitive intern compensation
  • Mentorship from researchers working on frontier voice AIEvery tool you need to succeed
  • Every tool you need to succeed
  • Beautiful office in Levi's Plaza, SF with rooftop views
  • A real shot at a return offer
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Intern
Machine Learning Intern

Bland AI • San Francisco (CA)

On-site
USD 40,000 - 65,000
Competitive intern compensation
Mentorship from researchers on front‑f
All the tools you need
+2
Audio ML Research Intern: Turn Research into Real Voice AI
Audio ML Research Intern: Turn Research into Real Voice AI

Bland AI • San Francisco (CA)

On-site
USD 40,000 - 65,000
Competitive intern compensation
Mentorship from researchers on front‑f
All the tools you need
+2
Machine Learning Research Intern 2027
Machine Learning Research Intern 2027

Mixpeek • San Francisco (CA), Northern (KY)

Hybrid
USD 34,000 - 48,000
Competitive internship pay
SF office in-person work
Meals provided in the office
Machine Learning Research Intern 2027
Machine Learning Research Intern 2027

Phonic • San Francisco (CA)

On-site
USD 45,000 - 65,000
Top-tier compensation
Free meals
Off-site & team events
Voice AI Research Intern: Build Real-World Speech Tech
Voice AI Research Intern: Build Real-World Speech Tech

Uncover • San Francisco (CA), Northern (KY)

Hybrid
USD 34,000 - 55,000
Mentorship from researchers
Office in Levi's Plaza, SF
Return offer potential
ML Researcher, Speech
ML Researcher, Speech

DeepRec.ai • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Healthcare
Dental
+1
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Machine Learning Researcher, Audio
Machine Learning Researcher, Audio

Bland AI • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Full healthcare, dental, vision
Meaningful equity
High autonomy, high impact
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Scientist - Audio [33340]
Research Scientist - Audio [33340]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000