Voice AI Engineer (Real-Time Speech)

N-iX

Kraków

Hybrid

PLN 240,000 - 360,000

Full time

20 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible remote/office option
Salary and benefits package
Career growth and trainings
Education reimbursement
Team events

Job summary

N-iX is seeking a Voice AI Engineer (Real-time speech) to design and operate end-to-end voicebot pipelines on AWS, integrating Whisper ASR and Azerbaijani TTS, with low latency and high throughput targets. You will work across telephony and chat platforms, ensure data privacy, and drive MLOps pipelines.

The role emphasizes GPU optimization, container orchestration, and secure, compliant data handling within a hybrid cloud setup. Flexible remote options are available.

Qualifications

  • 4+ years of hands-on experience with machine learning and Speech Processing with a primary focus on real-time conversational AI, ASR (STT), and TTS voice pipelines.
  • Deep expertise with Amazon SageMaker (real-time GPU inference endpoints, Pipelines, Feature Store, Model Registry) and Amazon Bedrock (AgentCore, Bedrock Guardrails, Knowledge Bases).
  • Proven track record in streaming speech inference, speech synthesis, and low-latency audio processing.
  • Strong experience in GPU optimization and containerized orchestration (NVIDIA A100/L40S, AWS EC2 GPU instances, Docker, Kubernetes/EKS).
  • Solid understanding of contact center and telephony platform integrations (Avaya, Genesys) and real-time decisioning interfaces.
  • Proficient in Python, Redis (priority queuing & caching), and data security/privacy (FPE tokenization, handling sensitive/PII data).

Responsibilities

  • Design, build, and operationalize end-to-end real-time STT → LLM → TTS voicebot pipelines on AWS, optimizing for streaming speech-to-text, first-token LLM generation, and first-audio TTS synthesis.
  • Deploy and maintain production customer-trained Whisper (Azerbaijani ASR) and Azerbaijani TTS models as low-latency real-time endpoints on Amazon SageMaker and specialized GPU node pools (NVIDIA A100/L40S).
  • Implement and manage the Bedrock Proxy Gateway on EKS for multi-model routing, priority queuing via Redis Sorted Sets, cost caps, and high-availability serving targeting ~200 rps without API throttling.
  • Integrate voicebot and chatbot decision engines with core enterprise telephony and CVM platforms, including Avaya (voice telephony), Genesys (digital chat/omnichannel), and Pelatro (CVM offer decisioning and uplift models).
  • Establish LLMOps & MLOps pipelines using Amazon SageMaker Pipelines and MLflow for experiment tracking, model versioning, prompt/agent registries, automated evaluation harnesses, and RAG knowledge base retrieval.
  • Build call and chat transcription pipelines to ingest, transcribe, and extract real-time insights (churn risk, dissatisfaction, intent, lead signals) into downstream decision layers.
  • Enforce data sovereignty and privacy controls by integrating on-premises Format Preserving Encryption (FPE) and tokenization wrappers into ML pipelines so zero raw PII enters AWS cloud environments.
  • Define NFR baselines, dialogue flows, voicebot persona, turn-taking, and fallback/escalation logic to guarantee conversational round-trip latency.
  • Automate ML deployment workflows using GitLab CI/CD and Infrastructure-as-Code (Terraform or AWS CDK), establishing observability and FinOps spend/anomaly monitoring via Amazon CloudWatch and Splunk.

Skills

Real-time
Speech processing
SageMaker
Bedrock
GPU optimization
Kubernetes/EKS
Python
Redis
Terraform/AWS CDK
GitLab CI/CD

Tools

AWS SageMaker
Amazon Bedrock
NVIDIA A100/L40S
Docker
Kubernetes

Job description

We are looking for a Voice AI Engineer (Real-time speech) to join our team!

Our client is an Azerbaijani telecommunications company and Azerbaijan's largest mobile network operator. The main products are: Fixed telephony, Mobile telephony, Internet services, Wireless broadband, and Value-added services. The primary goal is to accelerate the client’s Data & AI initiatives via a secure, hybrid cloud foundation on AWS while systematically modernizing the IT estate as part of the AWS MAP 2.0 program.

Key Project Objectives
  • Cloud Foundation & Landing Zone: Deploy target hybrid network architectures, establishing a secure AWS Landing Zone Accelerator (LZA) and hybrid Data/AI platforms on AWS.
  • Security, Compliance & Sovereignty: Operationalize on-premises data de-identification and Format Preserving Encryption (FPE) tokenization (achieving zero raw PII in the cloud), fully adhering to Azerbaijani Personal Data Law No. 998-IIIQ and Critical Information Infrastructure Rules (Resolution No. 229).
  • AI Chatbot & Real-Time Voicebot Implementation: Develop and operationalize a flagship Customer Care Voicebot (STT → LLM → TTS pipeline) and Agentic Chatbot targeting < 2.0s conversational voice latency and ~200 rps throughput as the first hybrid-setup consumer.
Responsibilities
  • Design, build, and operationalize end-to-end real-time STT → LLM → TTS (Speech-to-Text / LLM / Text-to-Speech) voicebot pipelines on AWS, optimizing for streaming speech-to-text, first-token LLM generation, and first-audio TTS synthesis.
  • Deploy and maintain production customer-trained Whisper (Azerbaijani ASR) and Azerbaijani TTS models as low-latency real-time endpoints on Amazon SageMaker and specialized GPU node pools (NVIDIA A100/L40S).
  • Implement and manage the Bedrock Proxy Gateway on EKS for multi-model routing, priority queuing via Redis Sorted Sets, cost caps, and high-availability serving targeting ~200 rps without API throttling.
  • Integrate voicebot and chatbot decision engines with core enterprise telephony and CVM platforms, including Avaya (voice telephony), Genesys (digital chat/omnichannel), and Pelatro (CVM offer decisioning and uplift models).
  • Establish LLMOps & MLOps pipelines using Amazon SageMaker Pipelines and MLflow for experiment tracking, model versioning, prompt/agent registries, automated evaluation harnesses, and RAG knowledge base retrieval.
  • Build call and chat transcription pipelines to ingest, transcribe, and extract real-time insights (churn risk, dissatisfaction, intent, lead signals) into downstream decision layers.
  • Enforce data sovereignty and privacy controls by integrating on-premises Format Preserving Encryption (FPE) and tokenization wrappers into ML pipelines so zero raw PII enters AWS cloud environments.
  • Define NFR baselines, dialogue flows, voicebot persona, turn-taking, and fallback/escalation logic to guarantee conversational round-trip latency.
  • Automate ML deployment workflows using GitLab CI/CD and Infrastructure-as-Code (Terraform or AWS CDK), establishing observability and FinOps spend/anomaly monitoring via Amazon CloudWatch and Splunk.
Requirements
  • 4+ years of hands-on experience with machine learning and Speech Processing with a primary focus on real-time conversational AI, ASR (STT), and TTS voice pipelines.
  • Deep expertise with Amazon SageMaker (real-time GPU inference endpoints, Pipelines, Feature Store, Model Registry) and Amazon Bedrock (AgentCore, Bedrock Guardrails, Knowledge Bases).
  • Proven track record in streaming speech inference, speech synthesis, and low-latency audio processing.
  • Strong experience in GPU optimization and containerized orchestration (NVIDIA A100/L40S, AWS EC2 GPU instances, Docker, Kubernetes/EKS).
  • Solid understanding of contact center and telephony platform integrations (Avaya, Genesys) and real-time decisioning interfaces.
  • Proficient in Python, Redis (priority queuing & caching), and data security/privacy (FPE tokenization, handling sensitive/PII data).
Nice-to-have skills
  • AWS Certified Machine Learning – Specialty or AWS Certified Solutions Architect.
  • Hands-on experience with EMR-on-EKS, Apache Iceberg, or MSK (Kafka) streaming pipelines.
Soft Skills & Team Fit
  • Strong critical thinking, problem-solving, and analytical skills with ownership of mission-critical, low-latency deliverables.
  • Excellent communication and collaboration skills to work closely with cross-functional teams (AI Architects, Data Engineers, CC SMEs, and Security/Compliance).
  • Results-oriented, proactive mindset with strong ownership within an Agile / Scrum framework.
  • Upper-Intermediate+ English level (written and spoken).
We offer
  • Flexible working format - remote, office-based or flexible
  • A competitive salary and good compensation package
  • Personalized career growth
  • Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
  • Active tech communities with regular knowledge sharing
  • Education reimbursement
  • Memorable anniversary presents
  • Corporate events and team buildings
  • Other location-specific benefits
  • not applicable for freelancers
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Agentic UX / Conversation Designer
Agentic UX / Conversation Designer

N-iX • Warszawa

Hybrid
PLN 120,000 - 180,000
Flexible working format
Professional development tools
Tech talks and trainings
Technical Program Manager (#5776)
Technical Program Manager (#5776)

N-iX • Poland

On-site
PLN 190,000 - 280,000
Remote Real-Time Voice AI Engineer—Whisper & LLM
Remote Real-Time Voice AI Engineer—Whisper & LLM

N-iX • Kraków

Hybrid
PLN 240,000 - 360,000
Flexible remote/office option
Salary and benefits package
Career growth and trainings
+2
Sr. AI Voice Engineer
Sr. AI Voice Engineer

Midway Auto Group • Polska

Hybrid
PLN 180,000 - 280,000
Senior DevOps Engineer/EKS
Senior DevOps Engineer/EKS

N-iX • Kraków

Hybrid
PLN 260,000 - 420,000
Flexible working format
Education reimbursement
Career growth
+4
Senior DataOps Engineer
Senior DataOps Engineer

N-iX • Warszawa

Hybrid
PLN 180,000 - 280,000
Flexible working format
Competitive salary
Career growth
+2
Senior DataOps Engineer
Senior DataOps Engineer

N-iX • Kraków

Hybrid
PLN 190,000 - 260,000
Flexible remote/office-based format
Competitive salary
Education reimbursement
+2
Senior AI Engineer
Senior AI Engineer

N-iX • Województwo małopolskie

On-site
PLN 120,000 - 150,000
Flexible working format
A competitive salary and good compensation package
Personalized career growth
+3
Senior AI Engineer
Senior AI Engineer

N-iX • Poland

On-site
PLN 80,000 - 110,000
Flexible working format
Competitive salary and good compensation package
Personalized career growth
+3
AI Architect (Voice AI)
AI Architect (Voice AI)

Neurons Lab • Warszawa

Hybrid
PLN 260,000 - 480,000