Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.
AGIGO in Switzerland invites a Master’s or PhD candidate (or recent graduate) to join a six-month internship in Voice & Conversational AI research. You will explore real-time audio profiling, evaluating encoder-based and micro speech-LLM approaches, and help ship a fast, reliable AudioProfile pipeline.
You will code in Python with PyTorch, collaborate with a team of researchers, and gain hands-on experience with ML operations, GPU acceleration, and deployment tooling.
STRUCTURED AUDIOPROFILING FOR CONVERSATIONAL AI AGENTS
Full-time | Voice & Conversational AI | Enterprise AI | Speech AI Team
Duration: 6 Months (flexible)
AGIGO provides the enterprise-grade conversational AI infrastructure and end-to-end toolchain to design and operate high-agency, human-like AI agents that engage directly with customers over phone, email, and text, handling complete customer interactions across support, bookings, and sales. AGIGO stands out by offering true AI sovereignty through on-premises deployment and zero exposure to third-party services. Powered by AGIGO’s proprietary technology stack, the platform delivers reliable agent operations, execution assurance, ultra-low latency, seamless enterprise integration, and predictable, token-free economics.
Founded in Switzerland in February 2025 by a team of experienced AI pioneers, AGIGO is building the infrastructure for a new generation of enterprise customer interactions, combining human-like communication with the control, reliability, and economics enterprises require at scale.
The moment a caller stops speaking, a voice agent has to answer a lot of questions about what it just heard. Which language and accent? Did they switch language mid-sentence, or finish their turn, or only pause to think? Did two people talk at once? Was there a number or a name in there? Is this a real voice or a synthetic one? Today each answer comes from a separate model, so every turn pays for a chain of forward passes: slow, expensive to run, and too big for the time budget of a live turn.
Build a highly optimized W2V2-sized model (or micro SpeechLLM) that runs in real time, and returns a highly structured and detailed AudioProfile that can help to understand the spoken audio message. Accuracy on any one task is not the hard part, since high-performance open models already exist (Whisper, PyAnnote, etc). The hard part is answering all of them on every turn inside a live call: eight specialists mean eight forward passes and eight deployments to maintain. The core question of this internship: does a frozen encoder with many small heads, or a micro speech-LLM that writes the AudioProfile as JSON and picks up new tasks from instructions is sufficient? You will build and compare both approaches and eventually help us in the deployment.
A frozen W2V2-sized encoder with one shared pass. These encoders hold different information at different depths, so you probe which layer feeds which head. Then the versioned AudioProfile schema, with a confidence per field and an explicit abstain state.
Language, accent and audio quality, trained from our 168k-hour, 77-language catalog using labels the annotation pipeline already produces. There is also space for synthetic data generation or bootstrapping already available data for training specific tasks.
Endpointing moves past silence thresholds to a voice-activity-projection objective, since cutting people off is the worst failure here. Then overlap, backchannel versus interruption, code-switch spans, content, identity, phonemes, spoofing and watermarking.
Compare both designs on the same tasks and serving harness. Encoder-and heads gives deterministic, quantizable output in one pass; a micro speech-LLM can take on tasks it was never trained for.
INT8 quantisation, ONNX Runtime with layer fusion, then batched serving, to reach the throughput target.
A registered, quantised model running live after each turn and offline over the catalog, where the same heads find the recordings we need instead of sampling and hoping.
As part of the internship, you will measure which tasks reinforce one another and which conflict
Short turns are hard for speaker models, and a caller moving from speakerphone to handset can look like a different person.
The intern will benchmark multiple approaches in order to find the ”sweet spot” of performance and latency.
The outcome of this work will replace a pipeline approach with a single highly optimized fast pass that says what is in the utterance and whether it is safe: routing to the right language and detecting code-switching, deciding whether the caller has finished, etc.
We value original thinking and encourage you to help shape and redefine the project’s direction as your research uncovers new insights. AGIGO fosters an open, collaborative environment where ideas can evolve freely. Exceptional innovation often emerges where disciplines and perspectives intersect, and we actively support creative exploration that pushes the boundaries of what Voice-AI can achieve.
[1] Layer-wise Analysis of a Self-supervised Speech Representation Model, 2021. https://arxiv.org/abs/2107.04734
[2] Voice Activity Projection: Self-supervised Learning of Turn-taking Events, 2022. https://arxiv.org/abs/2205.09812
[3] Qwen2-Audio Technical Report, 2024. https://arxiv.org/abs/2407.10759
AGIGO™ is a registered trademark of AGIGO AG, Switzerland.
By submitting your application, you agree to allow AGIGO to store and process your data for recruitment purposes. Unless otherwise requested, we may retain your data for up to one year to consider you for this or other future opportunities.