AI Architect (Voice AI)

Neurons Lab

Poland

Hybrid

PLN 300,000 - 420,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Neurons Lab is building a production-grade voice copilot for a leading US veterinary care network, taking a PoC into live production. You will own the full architecture—streaming STT, LLM field extraction, Chrome-extension delivery, and AWS infrastructure.

You will drive latency reductions, ensure strict data isolation, and run production tests with real metrics. You will mentor the AI Engineer and lead knowledge transfer, with a structured handover from the outgoing architect to ramp quickly.

Qualifications

  • Real-time voice pipelines and streaming STT experience.
  • LLM engineering including prompt design and extraction.
  • Observability, dashboards and cost/latency tracking.
  • Hands-on AWS Bedrock/serverless and SES knowledge.
  • Full-stack pragmatism with Python and some TS/Chrome extension.
  • Strong written and spoken English for US executives.

Responsibilities

  • Own the full pipeline from streaming STT to LLM extraction and extension delivery.
  • Drive latency improvements and reduce P95 latency.
  • Run model A/B tests with golden-set evaluation for accuracy.
  • Own evaluation, dashboards and per-call cost analysis and optimization.
  • Ensure production readiness with multi-concurrent calls and data isolation.
  • Deliver end-to-end epics with demo fallbacks and production polish.

Skills

Real-time voice pipelines
LLM engineering
Observability and evals
AWS Bedrock
Full-stack pragmatism
English communication

Tools

Langfuse
Chrome extension
AWS SES

Job description

About the project (description, duration, stage)

The client is the largest US network of in-home veterinary hospice and end-of-life care. A major US private-equity sponsor drives the AI program and plans more projects across its portfolio.

We built a real-time voice copilot for their Veterinary Care Coordinators (VCCs). The copilot listens to live calls with pet families. It extracts appointment and clinical fields while the call runs. It fills the client's scheduling system through a Chrome extension. A second workstream, the Vet Visit Copilot, sends each vet an AI pre-visit briefing by email (Amazon SES).

Next is the production phase.

Stage: production SOW in executive alignment; start expected September 2026.

Duration: multi-month, with strong extension probability. 0.5 FTE minimum; ramp toward 1.0 FTE as production scales.

Why the role is open: the current architect moves to another strategic build. He stays at 0.15-0.2 FTE for supervision and knowledge transfer during ramp-up, so the new architect gets a structured handover.

Objective
  • Own the technical architecture and delivery of the voice copilot from validated PoC to production

  • Hit the bar this client tests against: latency, accuracy, concurrency, and cost

  • Keep expectations aligned: production polish is in scope now; protect the team from silent scope creep

  • Transfer knowledge continuously to the client's team and Neurons Lab engineers

Areas of Responsibility
Technical architecture & hands-on implementation
  • Own the full pipeline: streaming speech-to-text, LLM field extraction, Chrome-extension delivery, and AWS infrastructure

  • Drive latency work: cut P95 from ~6s toward ~2s; remove post-processing corner cases (occasional ~1min lag on one field type)

  • Run model A/B tests (current pair: Claude Haiku vs GPT Luna) with golden-set evaluation for phonetic name and email accuracy

  • Own evaluation and cost: Langfuse traces, accuracy dashboards, real per-call cost from live calls, and an optimization plan

  • Harden for production: 5-10+ concurrent calls, strict data isolation between users, monitoring, alerting, and safe rollback

  • Ship epics end to end (example: the SES email briefing service); always keep a demo fallback so a live session never fails

Working with client stakeholders
  • Front technical discussions with a meticulous client; VCCs test edge cases and expect production quality

  • Present concrete system behavior, with numbers - this account rewards evidence, not slides

  • Hold the scope line: tie every feedback item to the SOW; route roadmap items (learning loop, persistent memory) to future phases

  • Keep internal discussions internal; all client-facing materials pass ADM review before sending

Team & knowledge
  • Lead the AI Engineer and the pod: set tasks, review output, unblock fast

  • Absorb the handover from the outgoing architect (0.15-0.2 FTE supervision window) and become independent fast

  • Run knowledge-transfer sessions; the project must have no single point of failure

  • Support the production SOW with estimates and architecture options when the account team asks

Skills
  • Real-time voice pipelines: streaming STT, turn handling, low-latency LLM inference - hands-on

  • LLM engineering: prompt engineering, structured extraction, guardrails, model A/B evaluation

  • Observability and evals: Langfuse or similar; golden datasets; latency, accuracy, and cost dashboards

  • AWS: Bedrock, serverless patterns, SES; token economics and per-call cost engineering

  • Full-stack pragmatism: strong Python; enough TypeScript / Chrome-extension knowledge to own the integration

  • Clear spoken and written English for demanding US executives

Knowledge
  • Contact-center / agent-assist patterns and metrics (handle time, cost per call, concurrency)

  • Production LLM operations: load testing, data isolation, incident handling

  • Nice to have: empathy-sensitive domains (healthcare, veterinary, insurance) and PE-sponsored rollouts

Experience

Key characteristics (screen for all four):

  1. Voice AI in production - mandatory. Shipped at least one real-time voice or speech product to real users (agent assist, voice bot, live transcription copilot). Candidates will demo real artifacts at the interview.

  2. 6+ years hands-on AI/ML engineering, with strong recent LLM production practice

  3. Latency and reliability record. Can show measured P95 reductions and concurrency fixes on a live system

  4. Consulting / client-facing seniority. Calm and precise under detailed UAT scrutiny; managers expectations well

Nice to have:

  • Chrome extension delivery; telephony / streaming stacks (Amazon Connect, Twilio, LiveKit)

  • Langfuse in production

  • US client experience with Eastern-time overlap

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Architect (Voice AI)
AI Architect (Voice AI)

Neurons Lab • Warszawa

Hybrid
PLN 260,000 - 480,000
Technical Lead
Technical Lead

SmartDev • Województwo kujawsko-pomorskie

Hybrid
PLN 420,000 - 640,000
Lead Voice AI Architect for Production Co-Pilot
Lead Voice AI Architect for Production Co-Pilot

Neurons Lab • Poland

Hybrid
PLN 300,000 - 420,000
Forward Deployed Engineer, Poland/CZ
Forward Deployed Engineer, Poland/CZ

Telnyx • Warszawa

Hybrid
PLN 180,000 - 320,000
Lead AI Engineer & Technical Architect (Remote)
Lead AI Engineer & Technical Architect (Remote)

Bold Business • Województwo małopolskie

On-site
PLN 254,000 - 382,000
Lead AI Engineer & Technical Architect (Remote)
Lead AI Engineer & Technical Architect (Remote)

Bold Business • Warszawa

On-site
PLN 80,000 - 110,000
AI-Native Engineer (Full-Stack / Agentic AI Engineer)
AI-Native Engineer (Full-Stack / Agentic AI Engineer)

SUNSCRAPERS Spółka z ograniczoną odpowiedzialnością • Warszawa

Hybrid
PLN 250,000 - 400,000
Cursor/Claude Pro licenses
Remote-first / flexible work
Direct interaction with CEO & VPs
Technical AI Engagement Lead
Technical AI Engagement Lead

Neurons Lab • Warszawa

On-site
PLN 300,000 - 420,000
AI-Native Engineer (Full-Stack / Agentic AI Engineer)
AI-Native Engineer (Full-Stack / Agentic AI Engineer)

Vecten • Warszawa

Hybrid
PLN 240,000 - 360,000
Lead/Staff Full Stack Engineer, AI Platform & Agents
Lead/Staff Full Stack Engineer, AI Platform & Agents

Wolters Kluwer • Województwo mazowieckie

On-site
PLN 212,000 - 319,000