Senior Machine Learning Engineer, Voice Agents

Jobgether

Deutschland

Vor Ort

EUR 90.000 - 140.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Remote work options
Health, dental, and vision benefits
Parental leave
Company equity
Education and Conference support
Office visits to NYC and Paris

Zusammenfassung

Jobgether, on behalf of a partner company, seeks a Senior Machine Learning Engineer for Voice Agents based in Germany. You will own architecture of the open-source speech-to-speech library, shaping APIs and streaming protocols while driving production-ready deployment in real-time systems.

The role blends backend and ML engineering with ASR/TTS integration, GPU serving, and developer experience. You will contribute to the community and help bring the platform to scale, including robotics

Qualifikationen

  • Senior-level engineering experience with the ability to independently own a substantial part of an architecture and drive it forward.
  • Experience building developer-facing infrastructure in AI, machine learning, developer tools, or a comparable technical environment, such as inference APIs or agent infrastructure.
  • Significant open-source contributions to a Python library and strong proficiency with asynchronous Python.
  • Solid understanding of distributed systems and their failure modes.
  • Proven experience shipping realtime technology involving streaming, WebSockets, WebRTC, audio or video pipelines, or live inference.
  • Practical production experience with LLMs or multimodal models.
  • Strong written communication skills and a demonstrated ability to collaborate asynchronously and in public.
  • Genuine interest in voice technology and conversational AI.
  • Contributions to voice-agent frameworks such as speech-to-speech, Pipecat, LiveKit Agents, Vocode, or TEN are a plus.

Aufgaben

  • Take architectural ownership of significant parts of the open-source speech-to-speech library, including pipeline design, latency budgets, and realtime loop reliability.
  • Integrate new ASR, TTS, and end-to-end speech models as they become available while maintaining clean, extensible abstractions.
  • Review community contributions, triage issues, manage releases, and help grow the contributor community around the project.
  • Design the developer API and streaming protocol for the voice platform, including session lifecycle, WebSockets/WebRTC transport, authentication, error semantics, and versioning.
  • Build and operate the serving infrastructure for realtime GPU inference, including concurrency, autoscaling, observability, and cost-per-session optimization.
  • Collaborate with Hub and inference teams to make voice agents easy to integrate into products, applications, and demonstrations.
  • Take the platform from prototype to production through load testing, SLO definition, reliability improvements, and graceful degradation when models or network paths fail.
  • Create documentation, examples, and templates that enable developers to move from initial setup to a running voice agent quickly.
  • Support deployments already relying on the technology, including the existing robotics fleet.
  • Contribute to the wider technical community through blog posts, demonstrations, conference talks, or other public technical content when desired.

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Engineer, Voice Agents based in Germany.

Own a major part of an open voice-agent stack at the intersection of machine learning, realtime systems, and developer infrastructure.
You will lead the architecture and evolution of an open-source speech-to-speech library while helping turn a new voice platform into a production-ready developer product.
The role combines deep backend and ML engineering with realtime audio, inference infrastructure, and developer experience.
You will integrate rapidly evolving ASR, TTS, and end-to-end speech models while maintaining clean abstractions and strong reliability.
You will have significant autonomy to shape APIs, streaming protocols, GPU serving, observability, and production architecture.
Your work will directly enable developers to build and deploy sophisticated voice agents and will support existing realtime deployments, including robotics applications.
This is an open, highly collaborative environment where you can contribute publicly through documentation, demos, talks, and open-source development.

Accountabilities
  • Take architectural ownership of significant parts of the open-source speech-to-speech library, including pipeline design, latency budgets, and realtime loop reliability.
  • Integrate new ASR, TTS, and end-to-end speech models as they become available while maintaining clean, extensible abstractions.
  • Review community contributions, triage issues, manage releases, and help grow the contributor community around the project.
  • Design the developer API and streaming protocol for the voice platform, including session lifecycle, WebSockets/WebRTC transport, authentication, error semantics, and versioning.
  • Build and operate the serving infrastructure for realtime GPU inference, including concurrency, autoscaling, observability, and cost-per-session optimization.
  • Collaborate with Hub and inference teams to make voice agents easy to integrate into products, applications, and demonstrations.
  • Take the platform from prototype to production through load testing, SLO definition, reliability improvements, and graceful degradation when models or network paths fail.
  • Create documentation, examples, and templates that enable developers to move from initial setup to a running voice agent quickly.
  • Support deployments already relying on the technology, including the existing robotics fleet.
  • Contribute to the wider technical community through blog posts, demonstrations, conference talks, or other public technical content when desired.
Requirements
  • Senior-level engineering experience with the ability to independently own a substantial part of an architecture and drive it forward.
  • Experience building developer-facing infrastructure in AI, machine learning, developer tools, or a comparable technical environment, such as inference APIs or agent infrastructure.
  • Significant open-source contributions to a Python library and strong proficiency with asynchronous Python.
  • Solid understanding of distributed systems and their failure modes.
  • Proven experience shipping realtime technology involving streaming, WebSockets, WebRTC, audio or video pipelines, or live inference.
  • Practical production experience with LLMs or multimodal models.
  • Strong written communication skills and a demonstrated ability to collaborate asynchronously and in public.
  • Genuine interest in voice technology and conversational AI.
  • Contributions to voice-agent frameworks such as speech-to-speech, Pipecat, LiveKit Agents, Vocode, or TEN are a plus.
  • Experience contributing to llama.cpp or another low-level inference runtime is advantageous.
  • Hands-on experience with ASR, TTS, or end-to-end speech models, including evaluating latency and quality trade-offs, is a plus.
  • GPU serving, quantization, or on-device inference experience is advantageous.
  • Knowledge of audio pipelines, including VAD, echo cancellation, jitter buffers, barge-in, and turn detection, is a plus.
  • Experience deploying technology to embedded or robotics environments is beneficial.
  • A public technical track record through talks, blog posts, demos, or similar contributions is valued.
  • Candidates are encouraged to apply even if they do not meet every listed requirement, particularly where their experience could bring complementary strengths to the team.
Benefits
  • Flexible working hours and remote work options.
  • Health, dental, and vision benefits for employees and their dependents.
  • Parental leave and flexible paid time off.
  • Company equity as part of the compensation package for all employees.
  • Reimbursement for relevant conferences, training, and education to support continuous professional development.
  • Access to a distributed, international work environment with opportunities to collaborate with experienced professionals across the AI and machine learning community.
  • Opportunities for remote employees to visit company offices in New York City and Paris.
  • Workstation equipment and setup support when needed to help employees work effectively.
  • A strong commitment to diversity, equity, and inclusion, with a workplace designed to ensure employees feel respected and supported.
  • Opportunities to contribute to and connect with the broader ML/AI community.
  • A culture focused on impact, continuous learning, collaboration, and professional growth.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Engineer – Agentic Systems (Voice) (m/f/d)
AI Engineer – Agentic Systems (Voice) (m/f/d)

DUDE CHEM • Berlin

Hybrid
EUR 90.000 - 130.000
Remote-friendly
Berlin office access
Hardware provided
AI Engineer – Agentic Systems (Voice) (m/f/d)
AI Engineer – Agentic Systems (Voice) (m/f/d)

Heyacto • Berlin

Hybrid
EUR 90.000 - 130.000
Remote-friendly with Berlin office
Flexible work schedule
30 days vacation
+2
Machine Learning Engineer
Machine Learning Engineer

ai|coustics • Berlin

Vor Ort
EUR 65.000 - 85.000
Competitive Compensation
Stock Options
Learning Opportunities
+4
Voice AI Engineer
Voice AI Engineer

PulseRise Technologies • Berlin

Vor Ort
EUR 85.000 - 160.000
Very competitive pay + equity package
Access to any AI tools
Free dinners if you stay late
+2
Engineering, Product Berlin
Engineering, Product Berlin

telli • Berlin

Vor Ort
EUR 90.000 - 130.000
Relocation support
Urban sports
Engineering, Voice AI Berlin
Engineering, Voice AI Berlin

telli • Berlin

Vor Ort
EUR 70.000 - 110.000
Relocation support if you don’t livein
Germany Senior / Lead Research Scientist - Germany
Germany Senior / Lead Research Scientist - Germany

Inworld AI • Deutschland

Vor Ort
USD 136.550 - 204.825
Engineering, All Berlin
Engineering, All Berlin

telli • Berlin

Vor Ort
EUR 90.000 - 130.000
Relocation support
Senior AI Platform Engineer with LangGrap
Senior AI Platform Engineer with LangGrap

Aether Biomedical • Deutschland

Hybrid
EUR 90.000 - 130.000
Vacation days
Illness days
Health insurance
+9
Systems Software Engineer (Rust, ML Inference)
Systems Software Engineer (Rust, ML Inference)

ai-coustics • Berlin

Vor Ort
EUR 60.000 - 80.000
Competitive salary package
Additional benefits and stock options
Dynamic startup culture