AI Research Engineer- Speech 1

Centific

Redmond (WA)

Hybrid

USD 150,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Benefits package
Hybrid/Remote options
GPU infrastructure access

Job summary

Centific, Redmond‑based AI research company, is seeking an AI Engineer: Speech/Audio to advance LALMs and S2S systems. You will work at the intersection of cutting‑edge research and production, designing spoken language models that understand, reason over, and generate audio with human‑like capabilities.

Join world‑class researchers and engineers, shape the technical direction in audio‑native AI, and contribute to publications; strong background in Python, PyTorch, and audio processing is

Qualifications

  • Master's degree in CS/EE or related field; PhD preferred.
  • 2+ years in speech/audio AI, LLMs, or multimodal systems.
  • Publications, patents, or shipped products in the field.

Responsibilities

  • Design and deploy Large Audio Language Models for native audio understanding and generation.
  • Develop Large Audio Reasoning Models with chain-of-thought reasoning for audio inputs.
  • Contribute to Speech-to-Speech systems including synthesis components.
  • Research alignment between speech encoders and LLM backbones using adapters and LoRA.
  • Optimize inference pipelines for low-latency, streaming speech applications.
  • Collaborate with teams to transfer research into production and publish results.

Skills

Python
PyTorch
Transformers
Speech processing
Multimodal learning
Research communication

Education

Master's degree in Computer Science or Electrical Engineering

Tools

GPU-accelerated training

Job description

About Centific

Centific is a frontier AI data foundry that curates diverse, high-quality data, using our purpose-built technology platforms to empower the Magnificent Seven and our enterprise clients with safe, scalable AI deployment. Our team includes more than 150 PhDs and data scientists, along with more than 4,000 AI practitioners and engineers. We harness the power of an integrated solution ecosystem—comprising industry-leading partnerships and 1.8 million vertical domain experts in more than 230 markets—to create contextual, multilingual, pre-trained datasets; fine-tuned, industry-specific LLMs; and RAG pipelines supported by vector databases. Our zero-distance innovation™ solutions for GenAI can reduce GenAI costs by up to 80% and bring solutions to market 50% faster. Our mission is to bridge the gap between AI creators and industry leaders by bringing best practices in GenAI to unicorn innovators and enterprise customers. We aim to help these organizations unlock significant business value by deploying GenAI at scale, helping to ensure they stay at the forefront of technological advancement and maintain a competitive edge in their respective markets.

About Job Job Description AI Engineer: Speech/Audio Centific AI Research
About Centific AI Research

Centific AI Research is at the forefront of developing cutting-edge AI solutions that bridge the gap between research innovation and real-world applications. Our team of scientists and engineers work collaboratively to create impactful technologies across speech, audio, and multimodal AI domains. We are committed to building responsible AI systems that deliver measurable impact while maintaining the highest standards of research quality.

The Opportunity

We are seeking an AI Engineer: Speech/Audio to join our growing team and drive innovation in next-generation audio AI technologies. This role focuses on Large Audio Language Models (LALMs), Large Audio Reasoning Models, and Speech-to-Speech (S2S) systems that can understand, reason over, and generate audio with human-like capabilities. You will work at the intersection of cutting-edge research and production systems, developing Spoken Language Models (SLMs) that perform complex audio reasoning and engage in natural speech-based interactions. This position offers the opportunity to shape our technical direction in audio-native AI while collaborating with world-class researchers and engineers.

Key Responsibilities
  • Design, develop, and deploy Large Audio Language Models (LALMs) capable of native audio understanding, reasoning, and generation.
  • Build Large Audio Reasoning Models that perform complex chain-of-thought reasoning over speech and audio inputs, including medical, technical, and conversational domains.
  • Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components.
  • Research and implement alignment mechanisms between speech encoders and LLM backbones using lightweight adapters, LoRA, and efficient fine-tuning strategies.
  • Design efficient speech tokenization and temporal compression techniques suitable for long-form audio reasoning and multi-turn spoken dialogue.
  • Build comprehensive evaluation frameworks for audio reasoning capabilities, including benchmarks for speech QA, audio understanding, and reasoning accuracy.
  • Optimize inference pipelines for low-latency, streaming applications in speech systems.
  • Collaborate with cross-functional teams to transfer research innovations into production systems and customer‑facing applications.
  • Contribute to technical documentation, research write-ups, and publications at top‑tier venues (NeurIPS, ICML, ACL, Interspeech).
Minimum Qualifications
  • Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning.
  • 2+ years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems.
  • Demonstrated applied research contributions through publications, patents, or shipped products in speech/audio AI or LLMs.
  • Strong proficiency in Python and PyTorch, with hands‑on experience in GPU‑accelerated training for large‑scale models.
  • Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations.
  • Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment methods.
  • Familiarity with modality alignment techniques: adapter‑based integration, cross‑modal attention, or audio‑text fusion methods.
  • Strong experimentation habits: clean code, systematic ablations, reproducibility, and clear technical communication.
Preferred Qualifications
  • Publication record at top‑tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP) in audio language models, speech reasoning, or multimodal learning.
  • Hands‑on experience building or fine‑tuning Large Audio Language Models (e.g., Qwen‑Audio, SALMONN, LTU, Gemini Audio).
  • Experience with speech representation pretraining (HuBERT, Wav2Vec 2.0, Whisper, WavLM) and discrete speech tokenization.
  • Familiarity with Speech-to‑Speech components: neural audio codecs (EnCodec, SoundStream), vocoders, or speech synthesis systems.
  • Experience with audio reasoning benchmarks (AIR‑Bench, MMAU, AudioBench) or building evaluation harnesses for audio QA.
  • Hands‑on experience with distributed training (FSDP, DeepSpeed) and inference optimization (ONNX, TensorRT, quantization).
  • Familiarity with speech frameworks such as ESPnet, SpeechBrain, NVIDIA NeMo, or Fairseq.
  • Experience with multilingual speech systems, code‑switching, or domain adaptation for specialized applications (medical, legal, technical).
  • Background in evaluating safety, bias, hallucination, or adversarial robustness in audio language models.
Technical Environment
  • Core: PyTorch, CUDA, torchaudio/librosa, Hugging Face Transformers
  • LLM Stack: Large language model backbones, lightweight adapters (LoRA, Q‑Former), instruction tuning pipelines
  • Audio Models: Neural audio codecs, speech encoders, vocoders, discrete speech tokenizers
  • >
What We Offer
  • Competitive compensation package with comprehensive benefits
  • Opportunity to work on cutting‑edge Large Audio Language Models and audio reasoning research with real‑world impact
  • Collaboration with experienced applied scientists and engineers in speech and multimodal AI
  • Support for publications at top‑tier conferences and professional development
  • Access to state‑of‑the‑art GPU infrastructure for training large‑scale audio models
  • Flexible work arrangements with hybrid/remote options

Location Redmond, WA / Palo Alto, CA / Remote Salary: $150-$160k Annually.

Centific is an equal‑opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status, or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.

Join a growing company using technology to help tackle enterprises’ toughest challenges.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Engineer- Speech
AI Research Engineer- Speech

Centific • Palo Alto (CA)

Hybrid
USD 150,000 - 210,000
Competitive pay
Hybrid work
Health coverage
+4
AI Research Engineer- Speech 1
AI Research Engineer- Speech 1

Centific Global Solutions, Inc. • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Competitive compensation
Hybrid/Remote options
GPU infrastructure access
+1
Senior AI Audio Scientist: LALMs & S2S (Remote)
Senior AI Audio Scientist: LALMs & S2S (Remote)

Centific • Redmond (WA)

Hybrid
USD 150,000 - 160,000
Benefits package
Hybrid/Remote options
GPU infrastructure access
Research, Audio Expertise
Research, Audio Expertise

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Audio Expertise
Research, Audio Expertise

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research, Audio Expertise
Research, Audio Expertise

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Sr Staff R&D Engineer
Sr Staff R&D Engineer

1421 Lucasfilm Ent Co Ltd, LLC Payroll Svc • California (MO)

On-site
USD 206,400 - 276,700
Medical benefits
Bonus
Principal Applied Scientist, Real-Time Conversational AI , AGI
Principal Applied Scientist, Real-Time Conversational AI , AGI

Amazon Science • Sunnyvale (CA)

On-site
USD 229,000 - 309,000
Health insurance
401(k) matching
Paid time off
+1
Senior Applied Scientist, Real-Time Conversational AI , AGI
Senior Applied Scientist, Real-Time Conversational AI , AGI

Amazon • Seattle (WA)

On-site
USD 167,000 - 226,000
Health insurance
401(k) matching
Paid time off
+1
Staff Research Scientist
Staff Research Scientist

techire ai • San Francisco (CA)

On-site
USD 350,000 - 400,000