AI Research Engineer- Speech 1

Centific Global Solutions, Inc.

Redmond (WA)

Hybrid

USD 150,000 - 160,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Hybrid/Remote options
GPU infrastructure access
Support for conferences/publications

Job summary

Centific Global Solutions, Inc. is seeking researchers to design and deploy Large Audio Language Models with native audio understanding and reasoning capabilities. You will build reasoning models for multi-domain speech inputs and contribute to S2S systems, including dialogue management and synthesis.

The role emphasizes research in alignment, tokenization, evaluation, and low-latency pipelines, with opportunities to publish and collaborate across teams.

Qualifications

  • Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning.
  • 2+ years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems.
  • Demonstrated applied research contributions through publications, patents, or shipped products in speech/audio AI or LLMs.
  • Strong proficiency in Python and PyTorch, with hands-on experience in GPU-accelerated training for large-scale models.
  • Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations.
  • Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment methods.
  • Familiarity with modality alignment techniques: adapter-based integration, cross-modal attention, or audio-text fusion methods.
  • Strong experimentation habits: clean code, systematic ablations, reproducibility, and clear technical communication.

Responsibilities

  • Design, develop, and deploy Large Audio Language Models (LALMs) capable of native audio understanding, reasoning, and generation.
  • Build Large Audio Reasoning Models that perform complex chain-of-thought reasoning over speech and audio inputs across medical, technical, and conversational domains.
  • Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components.
  • Research and implement alignment mechanisms between speech encoders and LLM backbones using lightweight adapters, LoRA, and efficient fine-tuning strategies.
  • Design efficient speech tokenization and temporal compression techniques suitable for long-form audio reasoning and multi-turn spoken dialogue.
  • Build comprehensive evaluation frameworks for audio reasoning capabilities, including benchmarks for speech QA, audio understanding, and reasoning accuracy.
  • Optimize inference pipelines for low-latency, streaming applications in speech systems.
  • Collaborate with cross-functional teams to transfer research innovations into production systems and customer-facing applications.
  • Contribute to technical documentation, research write-ups, and publications at top-tier venues (NeurIPS, ICML, ACL, Interspeech).

Skills

Python
PyTorch
GPU training
Speech/Audio ML
Transformers
Experimentation

Education

Master's degree in CS/EE or related field
Ph.D. preferred

Tools

ONNX
TensorRT
DeepSpeed
Weights & Biases

Job description

Key Responsibilities
  • Design, develop, and deploy Large Audio Language Models (LALMs) capable of native audio understanding, reasoning, and generation.
  • Build Large Audio Reasoning Models that perform complex chain-of-thought reasoning over speech and audio inputs across medical, technical, and conversational domains.
  • Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components.
  • Research and implement alignment mechanisms between speech encoders and LLM backbones using lightweight adapters, LoRA, and efficient fine-tuning strategies.
  • Design efficient speech tokenization and temporal compression techniques suitable for long‑form audio reasoning and multi‑turn spoken dialogue.
  • Build comprehensive evaluation frameworks for audio reasoning capabilities, including benchmarks for speech QA, audio understanding, and reasoning accuracy.
  • Optimize inference pipelines for low‑latency, streaming applications in speech systems.
  • Collaborate with cross‑functional teams to transfer research innovations into production systems and customer‑facing applications.
  • Contribute to technical documentation, research write‑ups, and publications at top‑tier venues (NeurIPS, ICML, ACL, Interspeech).
Minimum Qualifications
  • Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning.
  • 2+ years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems.
  • Demonstrated applied research contributions through publications, patents, or shipped products in speech/audio AI or LLMs.
  • Strong proficiency in Python and PyTorch, with hands‑on experience in GPU‑accelerated training for large‑scale models.
  • Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations.
  • Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment methods.
  • Familiarity with modality alignment techniques: adapter‑based integration, cross‑modal attention, or audio‑text fusion methods.
  • Strong experimentation habits: clean code, systematic ablations, reproducibility, and clear technical communication.
Preferred Qualifications
  • Publication record at top‑tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP) in audio language models, speech reasoning, or multimodal learning.
  • Hands‑on experience building or fine‑tuning Large Audio Language Models (e.g., Qwen‑Audio, SALMONN, LTU, Gemini Audio).
  • Experience with speech representation pretraining (HuBERT, Wav2Vec2.0, Whisper, WavLM) and discrete speech tokenization.
  • Familiarity with Speech‑to‑Speech components: neural audio codecs (EnCodec, SoundStream), vocoders, or speech synthesis systems.
  • Experience with audio reasoning benchmarks (AIR‑Bench, MMAU, AudioBench) or building evaluation harnesses for audio QA.
  • Hands‑on experience with distributed training (FSDP, DeepSpeed) and inference optimization (ONNX, TensorRT, quantization).
  • Familiarity with speech frameworks such as ESPnet, SpeechBrain, NVIDIA NeMo, or Fairseq.
  • Experience with multilingual speech systems, code‑switching, or domain adaptation for specialized applications (medical, legal, technical).
  • Background in evaluating safety, bias, hallucination, or adversarial robustness in audio language models.
Technical Environment
  • Core: PyTorch, CUDA, torchaudio/librosa, Hugging Face Transformers.
  • LLM Stack: Large language model backbones, lightweight adapters (LoRA, Q‑Former), instruction tuning pipelines.
  • Audio Models: Neural audio codecs, speech encoders, vocoders, discrete speech tokenizers.
  • Infrastructure: Modern GPU clusters, experiment tracking (Weights & Biases), distributed training frameworks.
  • Deployment: FastAPI/gRPC for services, ONNX/TensorRT for optimized inference.
What We Offer
  • Competitive compensation package with comprehensive benefits.
  • Opportunity to work on cutting‑edge Large Audio Language Models and audio reasoning research with real‑world impact.
  • Collaboration with experienced applied scientists and engineers in speech and multimodal AI.
  • Support for publications at top‑tier conferences and professional development.
  • Access to state‑of‑the‑art GPU infrastructure for training large‑scale audio models.
  • Flexible work arrangements with hybrid/remote options.

Location: Redmond, WA / Palo Alto, CA / Remote

Salary: $150‑$160k Annually.

Centific is an equal‑opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific • Redmond (WA)

On-site
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
ML Researcher, Speech
ML Researcher, Speech

DeepRec.ai • San Francisco (CA)

Hybrid
USD 250,000 - 300,000
Healthcare
Dental & Vision
Equity
+1
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+2
Member of Technical Staff - Multi-Modal, Audio San Francisco · Boston · Hybrid
Member of Technical Staff - Multi-Modal, Audio San Francisco · Boston · Hybrid

Liquid AI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
100% health premiums
401(k) matching up to 4%
Unlimited PTO
+1
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Research Scientist – Speech and Audio Understanding (Large Models & Multimodal Systems)
Research Scientist – Speech and Audio Understanding (Large Models & Multimodal Systems)

Lightspeed Studios • Bellevue (WA)

On-site
USD 123,000 - 230,000
Sign-on bonus
Relocation package
RSUs
Senior AI Audio Scientist: LALMs & S2S (Remote)
Senior AI Audio Scientist: LALMs & S2S (Remote)

Centific • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
Senior Inference Engineer, AGI
Senior Inference Engineer, AGI

Amazon • Sunnyvale (CA)

On-site
USD 192,000 - 260,000
Health insurance
401(k) matching
RSU/Stock options
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000
Equity
Member of Technical Staff - Multi-Modal, Audio
Member of Technical Staff - Multi-Modal, Audio

Liquid AI • San Francisco (CA)

Hybrid
USD 170,000 - 260,000
Equity
Health insurance
401(k) matching
+1