Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.
AGIGO invites Master’s or PhD students (preferred) or recent graduates in CS/ML to join a 6-month internship focused on low-latency, high-quality voice AI infrastructure.
You will build quantization pipelines, experiment with 4-bit weight representations, KF-cache quantisation, and sparsity while ensuring all variants meet a universal speech-quality gate across Hopper/Blackwell/ADA hardware.
Full-time | Voice & Conversational AI | Enterprise AI | Speech AI Team
Duration: 6 Months (flexible)
About AGIGO
AGIGO provides the enterprise-grade conversational AI infrastructure and end-to-end toolchain to design and operate high-agency, human-like AI agents that engage directly with customers over phone, email, and text, handling complete customer interactions across support, bookings, and sales. AGIGO stands out by offering true AI sovereignty through on-premises deployment and zero exposure to third-party services. Powered by AGIGO’s proprietary technology stack, the platform delivers reliable agent operations, execution assurance, ultra-low latency, seamless enterprise integration, and predictable, token-free economics.
Founded in Switzerland in February 2025 by a team of experienced AI pioneers, AGIGO is building the infrastructure for a new generation of enterprise customer interactions, combining human-like communication with the control, reliability, and economics enterprises require at scale.
Your Research Mission
Real-time voice agents require very low-latency decoding and increased serving costs compared to offline models. The latency becomes an extremely differentiating aspect, since a reply from an Voice Agent arriving after one second starts feeling broken. In this internship, you will build a complete pipeline for quantizaiton/quantization-aware training, and pruning for our internal LLMs and TTS models. Ideally, one initial checkpoint is compiled into a family of variants, each valid for a particular GPU generation (ADA/Hopper/Blackwell), precision, kernel stack, and batching regime, and every variant has to clear the same speech-aware quality gate before it is allowed out. Therefore, an important question arises: given the hardware and the latency requirements, which build are we allowed to serve? This matters because our stack is not uniform, but rather fluid, with Hopper, Blackwell or ADA machines requested on demand, which also might reward different recipes, so the same model has a different best answer depending on the initial conditions.
What You Will Build
Phase 1: Harness and gate
The benchmark harness and the CI quality gate, plus the schema for a build record: what it is valid for, and what it measured.
With and without in-domain calibration data.
Phase 3: Recovery
Quantisation-aware training and distillation, aiming to beat post-training quantisation at the same bit-width.
Phase 4: Sparsity
2:4 pruning, and investigate whether structured sparsity becomes real throughput.
Does the best build actually differ by hardware, and by how much? If both generations rank the variants the same way.
Is speculative decoding lossless for audio? You will investigate in which conditions speculative decoding for audio is lossless.
Your Impact
The resulting recipe of this internship will translate on optimizations in the models deployed in our stack, where each latency point and increase in tok/s really matters.
We value original thinking and encourage you to help shape and redefine the project’s direction as your research uncovers new insights. AGIGO fosters an open, collaborative environment where ideas can evolve freely. Exceptional innovation often emerges where disciplines and perspectives intersect, and we actively support creative exploration that pushes the boundaries of what Voice-AI can achieve.
What You Bring
Required
Bonus
What You Will Gain
Research in the Field
[1] AWQ: Activation-aware Weight Quantization for LLM Compression, MLSys 2024. https://arxiv.org/abs/2306.00978
[2] SmoothQuant: Post-Training Quantization for Large Language Models, 2022. https://arxiv.org/abs/2211.10438
[3] Principled Coarse-Grained Acceptance for Speculative Decoding in Speech, ICASSP 2026. https://arxiv.org/abs/2511.13732
AGIGO is a registered trademark of AGIGO AG, Switzerland.