Required Skill Set:* Experience fine-tuning open-source LLMs or SLMs (Small Language Models) specifically for high-speed conversational execution. * Prior experience building omnichannel systems where a single conversational engine powers both a text-based web/WhatsApp chatbot and a telephonic voice agent smoothly.
Experience: 1-2 Year
Job Summary:
We are looking for a Senior AI Engineer to spearhead the design, architecture, and deployment of our next-generation conversational AI ecosystems. In this role, you will bridge the gap between advanced Large Language Models (LLMs), Agentic workflows, and real-time voice infrastructure. You won't just be plugging into basic APIs; you will be optimizing low-latency audio pipelines, managing complex conversational state machines, and building intelligent agents that can handle unstructured, real-world human interactions seamlessly over both voice and chat interfaces.
Key Responsibilities:
Core Responsibilities
- End-to-End Voice Architecture: Design, build, and optimize scalable pipelines integrating Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and LLMs for ultra-low latency ( less than 1.5s perceived response time) voice bots.
- Agentic Framework Design: Develop and deploy multi-agent orchestration workflows that allow bots to reason, call external APIs, query databases dynamically, and execute complex business logic during live conversations.
- State & Context Management: Implement robust dialog management and context retention systems, ensuring the bot can handle interruptions, digressions, and multi-turn conversations without losing track.
- Backend & API Integration: Build highly performant backends to connect conversational interfaces with enterprise CRMs, databases, and third-party systems using robust API architectures.
- Evaluation & Optimization: Set up continuous evaluation frameworks to monitor bot performance, natural language understanding (NLU) accuracy, voice latency, and fallback rates, iteratively tuning prompts and fine-tuning models to improve user experience.
Technical Profile & Requirements
- Conversational Voice AI Expertise (Primary)
- Deep understanding of Telephony & WebRTC integration (e.g., Twilio, LiveKit, Vapi, Daily.co) for streaming audio.
- Hands-on experience optimizing ASR (Speech-to-Text) engines (Deepgram, Whisper, AssemblyAI) for real-time streaming, noise filtering, and accents.
- Experience with advanced TTS (Text-to-Speech) systems (ElevenLabs, Cartesia, OpenAI Voice) focusing on emotional inflection, prosody, and low-latency audio synthesis.
- Familiarity with handling voice-specific challenges like barge-in management (handling user interruptions mid-speech) and silence detection.
- Generative AI & Agentic Workflows
- Expertise in orchestrating LLMs (OpenAI GPT-4o, Claude 3.5 Sonnet, open-source Llama models) for conversational tasks.
- Proficiency in frameworks like LangChain, LangGraph, or CrewAI to build autonomous, tool-using AI agents.
- Advanced knowledge of RAG (Retrieval-Augmented Generation) techniques and Vector Databases (e.g., Pinecone, pgvector, Milvus) for grounding bot responses in specific datasets.
- Core Engineering Stack
- Languages: Strong mastery of Python (for AI framework integration and data processing).
- Backend Frameworks: Experience building asynchronous APIs using FastAPI or robust backend architectures using modern web frameworks.
- Databases: Proficiency in both relational databases ( MySQL, PostgreSQL ) and NoSQL systems for state persistence and logging.
- Streaming Protocols: Solid understanding of WebSockets and gRPC for bidirectional, real-time data streaming.
- Soft Skills & Best Practices
- Conversation Design Mindset: Ability to collaborate on structuring natural, empathetic, and efficient prompt designs that avoid robotic, cyclical loops.
- Production Discipline: Experience with CI/CD, containerization (Docker), logging/observability tools (LangSmith, Arize, or Prometheus), and managing AI infrastructure at scale.