Senior Machine Learning Engineer, Voice Agents - EMEA Remote

Hugging Face

Paris

Hybride

EUR 110 000 - 170 000

Plein temps

Il y a 10 jours
Générateur de candidature

Une candidature sur mesure pour ce poste — un CV personnalisé et une lettre de motivation qui correspondent directement à l’offre.

Passez les filtres ATS

Avantages offerts par ce poste

Remote options
Company equity
Conference reimbursement
Health benefits

Résumé du poste

Hugging Face is building a world-class open voice-agent stack and seeks a senior engineer to own a major portion of it. You will drive the speech-to-speech pipeline, integrate ASR/TTS models, and shape the hf-voice product while ensuring reliability and low latency.

You will review contributions, coordinate with hub and inference teams, and help take the product from demo to production with scalable serving and robust APIs.

Qualifications

  • Extensive open-source contributions to a Python library or similar.
  • Strong experience with async Python and distributed systems.
  • Ability to own architecture and drive it autonomously.
  • Experience shipping realtime streaming systems (WebSockets/WebRTC) and audio/video pipelines.
  • Familiarity with LLMs or multimodal models in production.

Responsabilités

  • Own the Open-Source Library: architecture design, latency budgeting, realtime loop reliability.
  • Integrate new ASR, TTS and end-to-end speech models as they land; maintain clean abstractions.
  • Review community PRs, triage issues, cut releases, grow contributors.

Connaissances

Open-source contributions
Async Python
Distributed systems
Latency optimization
Architectural ownership

Outils

WebSockets
WebRTC
GPU inference

Description du poste

At Hugging Face, we're on a journey to democratize good AI. We are building the fastest growing platform for AI builders with over 11 million users who collectively shared over 3M+ models, 1M+ datasets & 1.47M+ apps. Our open-source libraries have more than 600k+ stars on Github.

About the Role

We are building the open voice-agent stack for Hugging Face, and we are looking for a senior engineer to own a large part of it.Two things sit at the centre of this role. The first is speech-to-speech, our open-source library for realtime voice agents. The second is hf-voice, a new product that will let any developer build and deploy voice agents with their Hugging Face account.

The library already powers the Reachy Mini fleet and there is a public demo running on Spaces, so you won't start from a blank page. But almost everything about how this becomes a product developers rely on is still open, and you will have a direct say in it.

Your missions:
- Own the Open-Source Library:
  • Take architectural ownership of large parts of speech-to-speech: pipeline design, latency budget, and the reliability of the realtime loop
  • Integrate new ASR, TTS and end-to-end speech models as they land, and keep the abstractions clean while the model landscape keeps moving
  • Review community PRs, triage issues, cut releases, and grow the group of contributors around the project
- Ship hf-voice:
  • Design the developer API and the streaming protocol: session lifecycle, transport (WebSockets/WebRTC), authentication, error semantics, versioning
  • Build the serving side: realtime inference on GPU, concurrency, autoscaling, observability, and cost per session
  • Work with the Hub and inference teams so that a working voice agent is easy to integrate into products and demos
  • Take the product from demo to production: load testing, SLOs, graceful degradation when a model or a network path misbehaves
- Work in the Open:
  • Write the docs, examples and templates that get a developer from zero to a running agent in minutes
  • Support the deployments already relying on the stack, starting with the Reachy Mini fleet
  • Talk about the work publicly if you enjoy it: blog posts, demos, conference talks. We cover the travel and the prep time

At Hugging Face, we're on a journey to democratize good AI. We are building the fastest growing platform for AI builders with over 11 million users who collectively shared over 3M+ models, 1M+ datasets & 1.47M+ apps. Our open-source libraries have more than 600k+ stars on Github.

About the Role

We are building the open voice-agent stack for Hugging Face, and we are looking for a senior engineer to own a large part of it.Two things sit at the centre of this role. The first is speech-to-speech, our open-source library for realtime voice agents. The second is hf-voice, a new product that will let any developer build and deploy voice agents with their Hugging Face account.

The library already powers the Reachy Mini fleet and there is a public demo running on Spaces, so you won't start from a blank page. But almost everything about how this becomes a product developers rely on is still open, and you will have a direct say in it.

Your missions:
- Own the Open-Source Library:
  • Take architectural ownership of large parts of speech-to-speech: pipeline design, latency budget, and the reliability of the realtime loop
  • Integrate new ASR, TTS and end-to-end speech models as they land, and keep the abstractions clean while the model landscape keeps moving
  • Review community PRs, triage issues, cut releases, and grow the group of contributors around the project
- Ship hf-voice:
  • Design the developer API and the streaming protocol: session lifecycle, transport (WebSockets/WebRTC), authentication, error semantics, versioning
  • Build the serving side: realtime inference on GPU, concurrency, autoscaling, observability, and cost per session
  • Work with the Hub and inference teams so that a working voice agent is easy to integrate into products and demos
  • Take the product from demo to production: load testing, SLOs, graceful degradation when a model or a network path misbehaves
- Work in the Open:
  • Write the docs, examples and templates that get a developer from zero to a running agent in minutes
  • Support the deployments already relying on the stack, starting with the Reachy Mini fleet
  • Talk about the work publicly if you enjoy it: blog posts, demos, conference talks. We cover the travel and the prep time
Requirements
What we're looking for
  • Senior engineer, able to own a substantial part of an architecture and drive it forward autonomously
  • Experience building developer-facing infrastructure at an AI or developer-tools company: inference APIs, agent infrastructure, or something comparable
  • Substantial open-source contributions to a Python library. Comfortable with async Python and distributed systems, including their failure modes
  • You have shipped something realtime: streaming, WebSockets or WebRTC, audio or video pipelines, live inference
  • Practical experience with LLMs or multimodal models in production. Clear written communication and a habit of collaborating async and in public
  • Motivated by voice and conversational AI
Bonus points if you have
  • Contributions to a voice-agent framework such as speech-to-speech, pipecat, LiveKit Agents, Vocode or TEN
  • Contributions to llama.cpp or another low-level inference runtime
  • Hands-on work with ASR, TTS or end-to-end speech models, including evaluation of latency and quality trade-offs
  • GPU serving, quantization, or on-device inference experience
  • Audio pipeline knowledge: VAD, echo cancellation, jitter buffers, barge-in and turn detection
  • Experience shipping to embedded or robotics targets
  • A public track record: talks, blog posts, demos
About You

If you're interested in joining us, but don't tick every box above, we still encourage you to apply! We're building a diverse team whose skills, experiences, and backgrounds complement one another. We're happy to consider where you might be able to make the biggest impact.

One more thing

At Hugging Face we believe great AI shouldn't require a massive cluster, we build for everyone, especially the GPU-poor.

Benefits
More about Hugging Face

We are actively working to build a culture that values diversity, equity, and inclusivity. We are intentionally building a workplace where people feel respected and supported—regardless of who you are or where you come from. We believe this is foundational to building a great company and community. Hugging Face is an equal opportunity employer and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

We value development.You will work with some of the smartest people in our industry. We are an organization that has a bias for impact and is always challenging ourselves to continuously grow. We provide all employees with reimbursement for relevant conferences, training, and education.

We care about your well-being.We offer flexible working hours and remote options. We offer health, dental, and vision benefits for employees and their dependents. We also offer parental leave and flexible paid time off.

We support our employees wherever they are. While we have office spaces in NYC and Paris, we're very distributed and all remote employees have the opportunity to visit our offices. If needed, we'll also outfit your workstation to ensure you succeed.

We want our teammates to be shareholders. All employees have company equity as part of their compensation package. If we succeed in becoming a category-defining platform in machine learning and artificial intelligence, everyone enjoys the upside.

We support the community. We believe major scientific advancements are the result of collaboration across the field. Join a community supporting the ML/AI community.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior Open-Source Python Engineer, ML Developer Tools - EMEA Remote
Senior Open-Source Python Engineer, ML Developer Tools - EMEA Remote

Hugging Face • France

À distance
EUR 90 000 - 120 000
Company equity
Remote options
Health, dental, and vision benefits
+3
Cloud ML DevRel Engineer - EMEA remote
Cloud ML DevRel Engineer - EMEA remote

Hugging Face • Paris

Sur place
EUR 70 000 - 90 000
Flexible working hours
Health, dental, and vision benefits
Parental leave
+2
Low-level Senior Software Engineer, Xet Storage - EMEA Remote
Low-level Senior Software Engineer, Xet Storage - EMEA Remote

Hugging Face • Paris

Sur place
EUR 140 000 - 210 000
Remote options
Equity
Health benefits
+3
Senior ML Engineer, Open Voice Agents (Remote)
Senior ML Engineer, Open Voice Agents (Remote)

Hugging Face • Paris

Hybride
EUR 110 000 - 170 000
Remote options
Company equity
Conference reimbursement
+1
Platform Backend Engineer: APIs, TTS & AI (Async)
Platform Backend Engineer: APIs, TTS & AI (Async)

Speechify • Toulouse

Sur place
EUR 55 000 - 75 000
Competitive compensation
Dynamic work environment
Flexible work culture
Senior AI Software Engineer
Senior AI Software Engineer

GetVocal AI • Paris

Sur place
EUR 90 000 - 130 000
High ownership
Diverse international team
Cutting-edge AI exposure
+2
Software Engineer, Platform - Toulouse, France Toulouse, France
Software Engineer, Platform - Toulouse, France Toulouse, France

Speechify • Toulouse

Hybride
EUR 26 000 - 87 000
Senior AI Software Engineer
Senior AI Software Engineer

GetVocal AI Ltd. • Paris

Sur place
EUR 90 000 - 130 000
25 days holiday
Private healthcare
Diversified, international team
+1
Software Engineer, Platform - Lyon, France Lyon, France
Software Engineer, Platform - Lyon, France Lyon, France

Speechify • Lyon

Hybride
EUR 26 000 - 87 000
Senior Full-Stack AI Engineer (Incubator Team)
Senior Full-Stack AI Engineer (Incubator Team)

AI Digital • Job

Hybride
EUR 90 000 - 120 000