Software Engineer, Voice - Milan

indigo.ai

Milano

Remoto

EUR 40.000 - 70.000

Tempo pieno

3 giorni fa
Candidati tra i primi
Generatore di candidature

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Remote-friendly
Education budget
Meal vouchers
Top-grade equipment
Company retreats
Flexible work environment

Descrizione del lavoro

indigo.ai, a leading European AI scale-up, is seeking a Software Engineer, Voice to own the real-time voice stack for their AI agents. You’ll work on end-to-end voice pipelines—from telephony ingress to streaming STT/TTS—shaping how customers interact with AI on phone lines.

Ideal candidates have hands-on experience with live voice systems, latency budgets, and end-to-end production services using TypeScript/Node.js or Python, plus fluent English.

Competenze

  • Real-time voice or audio system design and shipping experience.
  • Experience with live telephony, streaming STT/TTS, and latency budgets.
  • Strong TypeScript/Node.js and/or Python for end-to-end services.
  • Experience with audio codecs and low-latency pipelines.
  • Fluent English communication.

Mansioni

  • Own the real-time voice pipeline end-to-end from ingress to TTS.
  • Engineer fast, human-like turn-taking and latency budgets.
  • Improve voice quality on the phone with STT/TTS evaluation.
  • Build and maintain production services with observability.
  • Evaluate and integrate STT/TTS providers and open-source tools.

Conoscenze

Real-time audio systems
Latency optimization
Voice UX
Production-grade software engineering
Fluent English

Strumenti

WebSockets/WebRTC/SIP
Opus codec
G.711/μ-law
Pipecat
LiveKit Agents
CI/CD
Observability tooling

Descrizione del lavoro

If you are here, it is because you know that we are looking for a Software Engineer, Voice for our Product team.

indigo.ai is the leading platform in Italy for building next-generation AI Agents that transform the way companies communicate with their customers. Since 2016, we’ve been helping enterprises in industries such as finance, insurance, utilities, retail, and e-commerce to evolve their Customer Experience through conversational AI. We don’t just “sell software”: we enable a shift in how organizations interact with people, automating millions of conversations every year, reducing operational costs, improving conversion rates, and creating more personalized, scalable, and compliant customer journeys.

Backed by a recent €10 million investment from Azimut, we are on a mission to take this technology global. This is a unique opportunity to join a well-funded, highly ambitious team and play a direct role in shaping the future of enterprise AI.

To make this happen, we are looking for a Software Engineer, Voice to support our Chief Product Development Officer in making talking to AI on the phone feel human.

What are we looking for?

We are looking for a Software Engineer, Voice to own the real-time voice layer of our AI Agents. Voice is where conversational AI is being decided right now, and our voice agents already handle production phone traffic for large companies. The bar is moving fast: we want to make the leap from "works reliably" to "feels human on a real phone line", and we want one person to own that leap. This is a specialist role with end-to-end ownership: the architecture, the model and provider choices, the latency budget, the way a conversation feels. You'll join our Product Engineering team, reporting to our Chief Product Development Officer, as our first full-time engineer dedicated to voice. And if voice grows the way we believe it will, you'll shape the team that grows around it.

Key Responsibilities:
  • Own the real-time voice pipeline end-to-end. From audio ingress on the telephony edge, through streaming STT and turn-taking, to the agent brain and back out through streaming TTS. Every millisecond in between is yours.

  • Engineer how fast the agent feels. Semantic end-of-turn detection, preemptive generation on partial transcripts, eager TTS, filler and backchannel strategies that mask tool calls. All measured on real 8kHz phone audio, not in a browser demo.

  • Make turn-taking human. Barge-in that survives noisy lines. Endpointing policies that know the dialog state, so a caller never gets cut off mid-IBAN. The difference between an IVR and a conversation lives here.

  • Raise voice quality on the channel that actually ships: the phone. Benchmark and A/B STT and TTS providers on real G.711 calls (Italian first: WER, naturalness, numbers and codes read right), exploit wideband/HD voice where the carrier allows it, and experiment with context-aware TTS and conversational speech models as they mature.

  • Build the evaluation harness. Turn "this voice sounds better" into numbers we trust: per-stage latency budgets, turn-taking metrics, regression suites on recorded calls, quality gates before anything reaches a client.

  • Keep production boringly reliable. Per-stage observability, live-call incident debugging (dead air, stuck turns, provider hiccups), graceful degradation when a vendor blinks.

  • Track a weekly-moving ecosystem and turn it into strategy. New STT/TTS/speech-to-speech releases land every month. You decide what we integrate, what we self-host for EU compliance and data residency, and what we skip. And you make provider swaps cheap.

You will need:

The filter is not your degree, and it's not years-of-experience arithmetic. It's having built it. Tell us about a real-time voice or audio system you designed and shipped: the latency budget, where it broke, and what you changed to make it feel right. That tells us more than any title.

  • Real-time audio systems, shipped. You've built voice agents, telephony systems, conferencing or live-streaming products that ran in production. You know what it means to move audio over WebSockets/WebRTC/SIP, through codecs (G.711/μ-law, Opus), against a latency budget.

  • The modern voice AI stack, hands-on. Streaming STT and TTS, VAD and turn detection, voice orchestration frameworks (Pipecat, LiveKit Agents or equivalent), speech-to-speech models. You have opinions on the trade-offs, grounded in things you've actually built, not blog posts.

  • Strong software engineering. TypeScript/Node.js and/or Python, and the maturity to own a production service end-to-end: containers, cloud infrastructure, CI/CD, observability.

  • A latency obsession. You think in milliseconds per stage, you instrument before you optimize, and you know the difference between measured and perceived latency, and how to exploit it.

  • A product ear. You can hear the difference between a demo and a conversation, and you can translate what you hear into engineering priorities and measurable evals.

  • An AI-native way of working. You use agentic coding tools (e.g. Claude Code) daily and you're good at directing them: setting up the problem, judging the output.

  • Language Skills: Fluent English.

We will really like (but they are not required):
  • Italian: our voice market is Italian-first, and you'll be tuning pronunciation, prosody and evals for it every week.

  • Contact-center / CCaaS ecosystem experience: SIP trunking, SBCs, enterprise telephony platforms.

  • ML audio experience: evaluating or fine-tuning ASR/TTS models, working with speech datasets.

  • Elixir: our agent platform is built on it.

  • Open source: contributions to open-source voice/audio projects.

Our values:

At indigo.ai we prize Vision (curiosity and courage to challenge the status quo), Connection (empathy, candor, and trust), Responsibility (ownership, reliability, and follow-through), and Excellence (the habit of raising the bar and refining until it’s right). If you naturally think ahead, build strong relationships, take accountability, and obsess over the quality of what you deliver, we’re probably a great match.

What do we offer?
  • A key role in one of Europe’s fastest-growing AI scale-ups, backed by a €10M investment from Azimut.

  • 40-70k RAL, commensurate with experience+ a performance-based Bonus.

  • A flexible, remote-friendly work environment.

  • Meal vouchers andWelfare programs to support your everyday life.

  • Access to a dedicated education budget for continued learning and growth.

  • Career Development Plan, ensuring a clear path for both personal and professional growth.

  • Top-grade equipment, which may include a MacBook Air, iPhone, and other top-tier devices.

  • Unlimited coffee.

  • Company retreats in stunning locations throughout the year.

Where is the job?

This position is fully remote, so there's no need to be in a specific location to do your work. That said, we have an amazing office at SPACES, Piazza Gae Aulenti 1/Torre B in Milan, available to anyone who wants to use it. We also love getting together and organize various retreats and meetups throughout the year to stay connected.

Why join us?

At indigo.ai you’ll be part of a fast-growing SaaS company where your ideas and workcan turn into products used by millions.Join us and you’ll:

  • Work with the most advanced AI technologies, shaping how enterprises across industries engage with their customers.

  • Be part of a dynamic and passionate team, where everyone has a direct impact on company growth.

  • Grow in a culture that values transparency, continuous learning, and career development, with clear paths for personal and professional progression.

  • Enjoy a flexible, people-first environment, with remote-friendly policies, stunning retreats, and a strong focus on well-being.

  • Contribute to our mission of reshaping how companies and people communicate worldwide.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Full-stack Developer, Senior
Full-stack Developer, Senior

indigo.ai • Milano

In loco
EUR 40.000 - 60.000
Competitive salary
Performance-based bonus
Flexible work environment
+4
AI Implementation Intern - Milan
AI Implementation Intern - Milan

indigo.ai • Milano

In loco
EUR 10.000 - 12.000
Hybrid work
Remote-friendly environment
Career development plan
+3
Full-stack Developer, Senior - Milan
Full-stack Developer, Senior - Milan

indigo.ai • Milano

In loco
EUR 70.000 - 100.000
Hybrid work model
Meal vouchers
Education budget
+4
Senior Financial Controller - Milan
Senior Financial Controller - Milan

indigo.ai • Turbigo

In loco
EUR 41.000 - 62.000
Competitive salary €41–62k RAL
Flexible, remote-friendly work environment
Meal vouchers
+4
Product Engineer, Agents
Product Engineer, Agents

Callimacus • Milano

In loco
EUR 45.000 - 65.000
Growth opportunities
Equity
Milan offices
+3
Strategic Account Executive - Italy
Strategic Account Executive - Italy

ElevenLabs • Milano

In loco
EUR 120.000 - 180.000
Annual discretionary stipend
Social travel stipend
Co‑working stipend
+1
Deployment Strategist - Italy
Deployment Strategist - Italy

Speedrun Talent Network • Milano

Ibrido
EUR 90.000 - 130.000
Strategic Account Executive - Italy Italy +1 more
Strategic Account Executive - Italy Italy +1 more

ElevenLabs • Milano

In loco
EUR 120.000 - 180.000
Innovative culture
Growth paths
Learning & development
+3
AI Engineer Senior - Fully Remote
AI Engineer Senior - Fully Remote

Aiclo • Lazio

Ibrido
EUR 70.000 - 110.000
Fully remote work
Autonomy in technical decisions
Direct collaboration with CTO
+3
Product Engineer, Platform
Product Engineer, Platform

Callimacus • Milano

In loco
EUR 45.000 - 65.000