Edge AI Researcher (Speech & Audio Models)

Huxley

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Huxley in San Francisco is seeking an Edge AI Research Scientist to advance speech and audio AI for devices with limited resources. You will design compact models, compress and optimize inference, and enable real-time speech experiences directly on phones and wearables.

You will work across ASR, TTS, speech translation, and neural audio codecs, combining research rigor with practical system engineering to deliver scalable, production-ready solutions.

Qualifications

  • Master's degree or Ph.D. in Computer Science or equivalent industry experience.
  • Strong expertise in model compression and edge AI for efficient inference.
  • Proficient in hardware-aware optimization and deployment on resource-constrained devices.
  • Fluent in PyTorch or JAX, and systems-level programming with C/C++.

Responsibilities

  • Drive research in efficient ML and edge AI for speech/audio on mobile devices.
  • Design compact architectures under strict latency, memory, and power constraints.
  • Develop and optimize training/inference techniques and compression methods.
  • Collaborate with ML researchers, mobile engineers, and systems teams to ship production-ready solutions.

Education

Master's degree in Computer Science or related field
Ph.D. or equivalent industry experience
Efficient Deep Learning

Tools

PyTorch
JAX
C/C++
CUDA
TensorRT
ONNX Runtime
Core ML
ExecuTorch
TensorFlow Lite
XNNPACK

Job description

We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI systems that run efficiently on smartphones, wearables, and other resource-constrained devices.

This role sits at the intersection of machine learning research, systems optimization, and production engineering. You will design compact model architectures, develop advanced compression techniques, and optimize inference pipelines that enable real-time speech AI experiences directly on end-user devices.

You will work across a broad range of voice technologies, including automatic speech recognition (ASR), text-to-speech (TTS), speech translation, speech-to-speech systems, and neural audio codecs.

The ideal candidate combines strong research credentials with hands-on implementation skills and has a deep understanding of efficient deep learning, model optimization, and hardware-aware machine learning.

What You'll Do

Research and Model Development

  • Drive research in efficient machine learning and edge AI for speech and audio applications.
  • Design compact model architectures capable of operating under strict latency, memory, and power constraints.
  • Develop and improve state-of-the-art approaches for:
  • Low-rank adaptation and compression
  • Hardware-aware architectures
  • Efficient training and inference techniques
  • Contribute to speech-to-speech, speech recognition, speech translation, text-to-speech, and audio generation systems.

Inference Optimization

  • Build highly optimized inference pipelines for mobile and embedded hardware.
  • Improve performance across CPUs, GPUs, NPUs, and other acceleration hardware.
  • Optimize:
  • Operator execution
  • Scheduling strategies
  • Caching mechanisms
  • End-to-end system latency
  • Integrate models with production runtimes and deployment frameworks.

Performance Evaluation

  • Develop rigorous benchmarking methodologies for edge AI systems.
  • Measure and improve:
  • Real-time factor
  • Time-to-first-audio
  • Thermal behavior
  • Speech quality and accuracy
  • Validate performance directly on target devices rather than relying solely on simulator environments.

Cross-Functional Collaboration

  • Partner with machine learning researchers, mobile engineers, and systems engineers to bring research into production.
  • Translate research prototypes into scalable products and customer-facing technologies.
  • Communicate findings through internal documentation, technical publications, conference papers, and open-source contributions where appropriate.

Required Qualifications

  • Master's degree, Ph.D., or equivalent industry experience in:
  • Computer Science
  • Efficient Deep Learning
  • Demonstrated expertise in model compression, efficient inference, or edge AI through research publications, production systems, or both.
  • Strong understanding of one or more of the following:
  • Hardware-aware optimization
  • Strong software engineering skills with:
  • PyTorch or JAX
  • C/C++ or equivalent systems-level programming experience
  • Experience optimizing neural networks for resource-constrained hardware.
  • Practical knowledge of:
  • GPUs
  • NPUs
  • Memory systems
  • Numerical precision tradeoffs
  • Ability to make informed tradeoffs between model quality, latency, memory footprint, power consumption, and deployment portability.
  • Professional proficiency in English.
  • Ability to thrive in a fast-moving, research-driven environment.

Preferred Qualifications

  • Experience working with:
  • Automatic Speech Recognition (ASR)
  • Speech-to-Speech Models
  • Familiarity with deployment frameworks such as:
  • Core ML
  • ExecuTorch
  • ONNX Runtime
  • LiteRT / TensorFlow Lite
  • TensorRT
  • Experience with acceleration technologies including:
  • Metal
  • Vulkan
  • CUDA
  • QNN
  • XNNPACK
  • Custom kernels and operator fusion
  • Knowledge of:
  • Experience building streaming and low-latency audio systems.
  • Experience deploying machine learning models to:
  • Embedded systems
  • Experience training, distilling, or evaluating large-scale foundation models using distributed GPU infrastructure.
  • Previous experience in industrial research labs, AI startups, or leading technology companies.
  • Publication record at relevant conferences or meaningful open-source contributions
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Machine Learning
Head of Machine Learning

5V Tech • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Bonus
Medical
+3
Embedded AI Engineer, On-Device Models
Embedded AI Engineer, On-Device Models

Deepgram • United States

On-site
USD 140,000 - 190,000
Edge AI Speech & Audio Researcher for Mobile
Edge AI Speech & Audio Researcher for Mobile

Huxley • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Scientist - Audio [33341]
Research Scientist - Audio [33341]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Research Scientist - Audio [33363]
Research Scientist - Audio [33363]

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 240,000
Tech Lead – ASR, TTS, Speech LLM, IC, Mentor
Tech Lead – ASR, TTS, Speech LLM, IC, Mentor

Jobtailor • Boston (MA)

On-site
USD 180,000 - 260,000
ML Researcher, Speech
ML Researcher, Speech

DeepRec.ai • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Healthcare
Dental
+1
Research Scientist - Speech
Research Scientist - Speech

JAM • United States

On-site
USD 100,000 - 130,000
Senior AI Engineer (Edge Dialog Systems)
Senior AI Engineer (Edge Dialog Systems)

Socket.dev • Palo Alto (CA)

Hybrid
USD 180,000 - 260,000
Audio AI Engineer (human)
Audio AI Engineer (human)

NEURA Robotics • Germany (OH)

On-site
USD 120,000 - 180,000