AI Architect (Voice AI)

Neurons Lab

Roma

In loco

EUR 90.000 - 130.000

Tempo pieno

5 giorni fa
Candidati tra i primi
Generatore di candidature

Ottieni una risposta da questo datore di lavoro — un curriculum e una lettera di presentazione personalizzati, che corrispondono esattamente a ciò che sta cercando.

Supera i filtri ATS

Descrizione del lavoro

Neurons Lab in Rome, Italy, is seeking a senior AI engineer to own the architecture and delivery of the voice copilot product for the veterinary care client. You will manage real-time streaming STT, LLM field extraction, Chrome-extension delivery, and the AWS Bedrock infrastructure to production readiness.

The role focuses on reducing latency, validating accuracy, and controlling per-call costs, with ongoing knowledge transfer to the client team and Neurons Lab engineers.

Competenze

  • Real-time voice AI in production — shipped at least one real-time voice or speech product to real users.
  • 6+ years hands-on AI/ML engineering with strong recent LLM production practice.
  • Latency improvements and concurrency fixes demonstrated on a live system.
  • Consulting / client-facing seniority; calm and precise under detailed UAT scrutiny.

Mansioni

  • Own the full pipeline: streaming STT, LLM field extraction, Chrome-extension delivery, and AWS infrastructure.
  • Drive latency work: reduce P95 from ~6s toward ~2s and remove post-processing lag.
  • Run model A/B tests with golden-set evaluation for name and email accuracy.
  • Own evaluation and cost: Langfuse traces, dashboards, and per-call cost optimization.
  • Harden for production: 5–10+ concurrent calls, data isolation, monitoring, and safe rollback.
  • Ship epics end to end (SES email briefing); maintain a demo fallback for reliability.

Conoscenze

Real-time voice pipelines
LLM engineering
Observability
Low-latency inference
Python
English communication
Cost optimization

Strumenti

Langfuse
AWS Bedrock
SES
Chrome Extension

Descrizione del lavoro

About The Project (description, Duration, Stage)

The client is the largest US network of in-home veterinary hospice and end-of-life care. A major US private-equity sponsor drives the AI program and plans more projects across its portfolio.

We built a real-time voice copilot for their Veterinary Care Coordinators (VCCs). The copilot listens to live calls with pet families. It extracts appointment and clinical fields while the call runs. It fills the client's scheduling system through a Chrome extension. A second workstream, the Vet Visit Copilot, sends each vet an AI pre-visit briefing by email (Amazon SES).

Next is the production phase.

Stage: production SOW in executive alignment; start expected September 2026.

Duration: multi-month, with strong extension probability. 0.5 FTE minimum; ramp toward 1.0 FTE as production scales.

Why the role is open: the current architect moves to another strategic build. He stays at 0.15–0.2 FTE for supervision and knowledge transfer during ramp-up, so the new architect gets a structured handover.

Objective
  • Own the technical architecture and delivery of the voice copilot from validated PoC to production
  • Hit the bar this client tests against: latency, accuracy, concurrency, and cost
  • Keep expectations aligned: production polish is in scope now; protect the team from silent scope creep
  • Transfer knowledge continuously to the client's team and Neurons Lab engineers
Areas of Responsibility
Technical architecture & hands-on implementation
  • Own the full pipeline: streaming speech-to-text, LLM field extraction, Chrome-extension delivery, and AWS infrastructure
  • Drive latency work: cut P95 from ~6s toward ~2s; remove post-processing corner cases (occasional ~1min lag on one field type)
  • Run model A/B tests (current pair: Claude Haiku vs GPT Luna) with golden-set evaluation for phonetic name and email accuracy
  • Own evaluation and cost: Langfuse traces, accuracy dashboards, real per-call cost from live calls, and an optimization plan
  • Harden for production: 5–10+ concurrent calls, strict data isolation between users, monitoring, alerting, and safe rollback
  • Ship epics end to end (example: the SES email briefing service); always keep a demo fallback so a live session never fails
Working with client stakeholders
  • Front technical discussions with a meticulous client; VCCs test edge cases and expect production quality
  • Present concrete system behavior, with numbers — this account rewards evidence, not slides
  • Hold the scope line: tie every feedback item to the SOW; route roadmap items (learning loop, persistent memory) to future phases
  • Keep internal discussions internal; all client-facing materials pass ADM review before sending
Team & knowledge
  • Lead the AI Engineer and the pod: set tasks, review output, unblock fast
  • Absorb the handover from the outgoing architect (0.15–0.2 FTE supervision window) and become independent fast
  • Run knowledge-transfer sessions; the project must have no single point of failure
  • Support the production SOW with estimates and architecture options when the account team asks
Skills
  • Real-time voice pipelines: streaming STT, turn handling, low-latency LLM inference — hands-on
  • LLM engineering: prompt engineering, structured extraction, guardrails, model A/B evaluation
  • Observability and evals: Langfuse or similar; golden datasets; latency, accuracy, and cost dashboards
  • AWS: Bedrock, serverless patterns, SES; token economics and per-call cost engineering
  • Full-stack pragmatism: strong Python; enough TypeScript / Chrome-extension knowledge to own the integration
  • Clear spoken and written English for demanding US executives
Knowledge
  • Contact-center / agent-assist patterns and metrics (handle time, cost per call, concurrency)
  • Production LLM operations: load testing, data isolation, incident handling
  • Nice to have: empathy-sensitive domains (healthcare, veterinary, insurance) and PE-sponsored rollouts
Experience

Key characteristics (screen for all four):

  • Voice AI in production — mandatory. Shipped at least one real-time voice or speech product to real users (agent assist, voice bot, live transcription copilot). Candidates will demo real artifacts at the interview.
  • 6+ years hands-on AI/ML engineering, with strong recent LLM production practice
  • Latency and reliability record. Can show measured P95 reductions and concurrency fixes on a live system
  • Consulting / client-facing seniority. Calm and precise under detailed UAT scrutiny; manages expectations well
Nice to have:
  • Chrome extension delivery; telephony / streaming stacks (Amazon Connect, Twilio, LiveKit)
  • Langfuse in production
  • US client experience with Eastern-time overlap
Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

AI Engineer
AI Engineer

Expert System S.p.A • Modena

Ibrido
EUR 45.000 - 65.000
Production AI Architect | Real-Time Voice Copilot Lead
Production AI Architect | Real-Time Voice Copilot Lead

Neurons Lab • Roma

In loco
EUR 90.000 - 130.000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Skillvue • Milano

Remoto
EUR 70.000 - 90.000
Competitive compensation
Flexible work
Budget for conferences and training
+1
Forward Deployed Engineer, Southern Europe
Forward Deployed Engineer, Southern Europe

Telnyx • Milano

Ibrido
EUR 90.000 - 120.000
Product Engineer, Agents
Product Engineer, Agents

Callimacus • Milano

In loco
EUR 45.000 - 65.000
Growth opportunities
Equity
Milan offices
+3
AI Engineer
AI Engineer

NTT DATA Europe & Latam • Ro

In loco
EUR 75.000 - 110.000
AI/ML Team Lead – Generative AI (LLMs, AWS)
AI/ML Team Lead – Generative AI (LLMs, AWS)

Provectus • Italia

Remoto
EUR 102.000 - 120.000
Sign-up bonus 8,000 USD
Long-term B2B collaboration
Fully remote setup
+3
Product Engineer, Agents
Product Engineer, Agents

Callimacus • Lazio

In loco
EUR 45.000 - 55.000
Equity
Uffici centrali a Milano
Assicurazione sanitaria
+2
Product Engineer, Platform
Product Engineer, Platform

Callimacus • Milano

In loco
EUR 45.000 - 65.000
Head of AI-Native Product Operations
Head of AI-Native Product Operations

Lansweeper • Roma

In loco
EUR 45.000 - 70.000