Research Internship: Efficient Inference for Text-to-Speech Models and LLMs

AGIGO

Zürich

Vor Ort

CHF 20.000 - 29.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Paid internship
Flexible hours
Access to GPU clusters (Hopper, Blackw

Zusammenfassung

AGIGO invites Master’s or PhD students (preferred) or recent graduates in CS/ML to join a 6-month internship focused on low-latency, high-quality voice AI infrastructure.

You will build quantization pipelines, experiment with 4-bit weight representations, KF-cache quantisation, and sparsity while ensuring all variants meet a universal speech-quality gate across Hopper/Blackwell/ADA hardware.

Qualifikationen

  • Master’s or PhD student (preferred) or recent graduate in Computer Science, Machine Learning, or a related field
  • Strong Python programming skills and Git
  • Solid understanding of ML fundamentals and MLOps
  • Hands-on experience with PyTorch
  • Fluent in English, highly motivated, willingness to learn

Aufgaben

  • Develop quantization-aware training pipelines for internal LLMs and TTS models
  • Implement quantization, pruning, and hardware-adaptive build variants
  • Build and maintain a quality gate with metrics like UTMOS, WER, and speaker similarity
  • Explore elastic serving and speculative decoding for multiple hardware targets
  • Collaborate with researchers and engineers to translate research into deployable optimizations

Kenntnisse

Python programming
Git
ML fundamentals
PyTorch
English fluency

Ausbildung

Master’s or PhD student (preferred) in CS/ML
Recent graduate in CS/ML

Tools

N/A

Jobbeschreibung

Full-time | Voice & Conversational AI | Enterprise AI | Speech AI Team

Duration: 6 Months (flexible)

About AGIGO

AGIGO provides the enterprise-grade conversational AI infrastructure and end-to-end toolchain to design and operate high-agency, human-like AI agents that engage directly with customers over phone, email, and text, handling complete customer interactions across support, bookings, and sales. AGIGO stands out by offering true AI sovereignty through on-premises deployment and zero exposure to third-party services. Powered by AGIGO’s proprietary technology stack, the platform delivers reliable agent operations, execution assurance, ultra-low latency, seamless enterprise integration, and predictable, token-free economics.

Founded in Switzerland in February 2025 by a team of experienced AI pioneers, AGIGO is building the infrastructure for a new generation of enterprise customer interactions, combining human-like communication with the control, reliability, and economics enterprises require at scale.

Your Research Mission

Real-time voice agents require very low-latency decoding and increased serving costs compared to offline models. The latency becomes an extremely differentiating aspect, since a reply from an Voice Agent arriving after one second starts feeling broken. In this internship, you will build a complete pipeline for quantizaiton/quantization-aware training, and pruning for our internal LLMs and TTS models. Ideally, one initial checkpoint is compiled into a family of variants, each valid for a particular GPU generation (ADA/Hopper/Blackwell), precision, kernel stack, and batching regime, and every variant has to clear the same speech-aware quality gate before it is allowed out. Therefore, an important question arises: given the hardware and the latency requirements, which build are we allowed to serve? This matters because our stack is not uniform, but rather fluid, with Hopper, Blackwell or ADA machines requested on demand, which also might reward different recipes, so the same model has a different best answer depending on the initial conditions.

What You Will Build

  • The build matrix. An initial checkpoint in, let’s say FP16/BF16 precision, then build a matrix from: weight-only 4-bit, activation quantisation with outlier handling, FP8 and NVFP4, KV-cache quantisation, and 2:4 sparsity where the hardware can use it. Each build is tagged with the hardware and workload it is valid for.
  • The quality gate. You will implement strong evaluation pipelines to signal issues in a quantized checkpoint, e.g., UTMOS, word error rate or speaker similarity for TTS, or other metrics such as intent accuracy or NER performance for LLMs.
  • Elastic serving and speculative decoding. One checkpoint offering several operating points chosen per call rather than per deployment; and speculative decoding for specific LLMs tasks or TTS.

Phase 1: Harness and gate

The benchmark harness and the CI quality gate, plus the schema for a build record: what it is valid for, and what it measured.

With and without in-domain calibration data.

Phase 3: Recovery

Quantisation-aware training and distillation, aiming to beat post-training quantisation at the same bit-width.

Phase 4: Sparsity

2:4 pruning, and investigate whether structured sparsity becomes real throughput.

Does the best build actually differ by hardware, and by how much? If both generations rank the variants the same way.

Is speculative decoding lossless for audio? You will investigate in which conditions speculative decoding for audio is lossless.

Your Impact

The resulting recipe of this internship will translate on optimizations in the models deployed in our stack, where each latency point and increase in tok/s really matters.

We value original thinking and encourage you to help shape and redefine the project’s direction as your research uncovers new insights. AGIGO fosters an open, collaborative environment where ideas can evolve freely. Exceptional innovation often emerges where disciplines and perspectives intersect, and we actively support creative exploration that pushes the boundaries of what Voice-AI can achieve.

What You Bring

Required

  • Current Master’s or PhD student (preferred), or recent graduate in Computer Science, Machine Learning, or a related degree field
  • Strong Python programming skills and Git
  • Solid understanding of ML fundamentals and MLOps
  • Hands-on experience with PyTorch
  • Fluent in English, highly motivated, willingness to learn

Bonus

  • CUDA, mixed precision, profiling, inference servers, or quantization tools such as: llm-compressor (vLLM) or Model-Optimizer (NVIDIA).

What You Will Gain

  • Direct impact on our product: your code ships in our platform, built alongside our researchers and engineers
  • Mentorship: work closely with our expert team of researchers and engineersTop-tier AI infrastructure: access to GPU clusters with NVIDIA Hopper (H200) and Blackwell RTX 6000 PRO NVIDIA GPUs
  • Research visibility: we will actively support you in publishing your work at a top-tier conference or in a journal paper
  • Disciplined and inspiring research environment: a team of sharp minds grounded in expertise, autonomy, and a shared pursuit of impactful breakthroughs
  • Paid internship: market-level salary, flexible hours, unlimited coffee, drinks, fruit and snacks
  • Career path: this internship may lead to a full-time permanent role in AGIGO's world-class AI R&D team

Research in the Field

[1] AWQ: Activation-aware Weight Quantization for LLM Compression, MLSys 2024. https://arxiv.org/abs/2306.00978

[2] SmoothQuant: Post-Training Quantization for Large Language Models, 2022. https://arxiv.org/abs/2211.10438

[3] Principled Coarse-Grained Acceptance for Speculative Decoding in Speech, ICASSP 2026. https://arxiv.org/abs/2511.13732

AGIGO is a registered trademark of AGIGO AG, Switzerland.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Research Internship: Audio Toolbox and (Micro) Speech-LLMs
Research Internship: Audio Toolbox and (Micro) Speech-LLMs

AGIGO • Zürich

Vor Ort
CHF 28.000 - 39.000
Paid internship
Market-level salary
Flexible hours
+1
Research Intern: Low-Latency Speech & LLM Inference
Research Intern: Low-Latency Speech & LLM Inference

AGIGO • Zürich

Vor Ort
CHF 20.000 - 29.000
Paid internship
Flexible hours
Access to GPU clusters (Hopper, Blackw
Paid AI Research Intern - Flexible Hours & Perks
Paid AI Research Intern - Flexible Hours & Perks

AGIGO • Zürich

Vor Ort
CHF 28.000 - 39.000
Paid internship
Market-level salary
Flexible hours
+1
Open Audio ML Research (Master Thesis / Internship)
Open Audio ML Research (Master Thesis / Internship)

Embodied AI • Zürich

Hybrid
CHF 20.000 - 27.000
Paid Master thesis
Own Topic
Real Attack Data
+2
ML / Data Engineer (Master Thesis / Internship)
ML / Data Engineer (Master Thesis / Internship)

Embodied AI • Zürich

Hybrid
CHF 20.000 - 36.000
Paid Master thesis or internship
One problem that is genuinely yours
Your work goes into the training stack
+1
AI Research Scientist
AI Research Scientist

Giotto.ai • Lausanne

Hybrid
CHF 140.000 - 210.000
Lead Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems
Lead Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems

Giotto.ai SA • Lausanne

Hybrid
CHF 140.000 - 190.000
AI Research Scientist
AI Research Scientist

Embodied AI • Lausanne

Hybrid
CHF 120.000 - 170.000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Schweiz

Vor Ort
CHF 100.000 - 140.000
Member of Technical Staff, NeuroAI
Member of Technical Staff, NeuroAI

Neurosoft Bioelectronics • Genf

Hybrid
CHF 45.000 - 55.000
Hybrid work model
Office in Geneva
Equity participation may be available