Lead AI Engineer Agentic Systems And Voice AI

YAL

Hyderabad

On-site

INR 3,000,000 - 6,000,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

YAL seeks a Lead AI Engineer to architect and deploy ultra-low-latency, agentic voice systems for Indian languages. You will bridge client workflows with scalable AI, ensuring robust backends and real-time observability at enterprise scale.

Responsibilities include designing cascaded ASR-LLM-TTS pipelines, telephony API integration, and optimizing latency metrics like TTFT and RTF while handling massive production traffic.

Qualifications

  • 8+ years overall software/AI/ML experience required.
  • Minimum 3+ years architecting agentic LLM systems in production.
  • Strong backend, distributed systems, and high-concurrency expertise.
  • Experience with Kafka, Temporal, and large-scale databases.

Responsibilities

  • Design and deploy ultra-low-latency conversational voice pipelines.
  • Integrate telephony APIs with AI backends for inbound/outbound calls.
  • Lead multilingual Indian language support and code-switching challenges.
  • Optimize TTFT/TTFA/RTF and model serving for high load.

Skills

8+ years software/AI/ML experience
3+ years architecting agentic LLMs
Backend for high-concurrency
Distributed systems
Voice/ASR/TTS knowledge
High-performance inference serving

Tools

Apache Kafka
Temporal
PostgreSQL
ClickHouse
WebRTC
gRPC
WebSockets
NVIDIA Triton
TensorRT
PyTorch
vLLM
HuggingFace
Twilio/Exotel

Job description

Job Description:
Role Overview

We are seeking a highly experienced Lead AI Engineer to architect, deploy, and scale intelligent, agentic voice systems capable of handling massive production traffic. This role focuses on building ultra-low-latency, full-duplex conversational voice agents using cascaded architectures (ASR -> LLM -> TTS) tailored for Indian languages. You will act as the technical bridge between complex client workflows and scalable AI capabilities, ensuring robust backend execution, optimal resource efficiency, and seamless human-computer interaction at enterprise scale.

Core Responsibilities
  • Full-Duplex Voice Architecture: Design and orchestrate conversational AI pipelines without relying on native multimodal/full-duplex LLMs. Build and tune highly responsive turn-taking logic, Voice Activity Detection (VAD), and barge-in/interruption handling across cascaded ASR, LLM, and TTS components.
  • Telephony & API Integration: Seamlessly bridge AI inference pipelines with standard telephony APIs (Twilio, Plivo, Exotel) for inbound and outbound agent call flows, managing basic call state (transfers, hold, drop detection).
  • Agentic Workflow Orchestration: Collaborate directly with clients to deconstruct complex business requirements and operational workflows. Translate these into deterministic agentic capabilities, utilizing state machines, tool-calling, and external API integrations to execute multi-step reasoning tasks.
  • Latency & Efficiency Optimization: Drive hardcore performance tuning across the entire stack. Optimize Time-To-First-Token (TTFT), Time-To-First-Audio (TTFA), and Real-Time Factor (RTF) over phone lines. Implement model quantization, KV cache optimization, dynamic batching, and efficient model serving to minimize latency under heavy concurrent loads.
  • Indic Language Mastery: Lead the development of multilingual systems that natively handle the phonetic and linguistic nuances of Indian languages. Solve complex challenges related to code-switching (e.g., Hindi-English, Telugu-English), regional accents, and low-resource language modeling.
  • Backend & Production Scale: Architect resilient, event-driven backend systems capable of sustaining high-throughput production traffic. Manage stateful asynchronous processes, distributed microservices, and robust data pipelines to ensure zero-downtime deployments and real-time observability.
Required Qualifications & Experience
  • Experience Baseline: 8+ years of overall software engineering and AI/ML experience, with a strict minimum of 3+ years directly architecting and deploying agentic LLM systems and complex conversational AI in production.
  • Production System Expertise: Deep understanding of backend engineering for high-concurrency environments. Proven experience with distributed systems, event-driven architectures (e.g., Apache Kafka), workflow orchestrators (e.g., Temporal), and high-performance databases (e.g., PostgreSQL, ClickHouse).
  • Conversational AI Depth: Strong operational knowledge of speech processing models (ASR/TTS) and streaming protocols (WebRTC, gRPC, WebSockets). You must know how to handle endpointing, stream buffering, and state management for natural voice interactions.
  • Optimization & Serving: Hands-on experience with high-performance inference servers (e.g., vLLM, NVIDIA Triton, TensorRT) and optimization techniques for large-scale model deployment.
  • Client to Code Translation: Demonstrated ability to act as a technical architect who can sit with stakeholders, map out domain-specific workflows (e.g., public grievance handling, CRM automation), and model them into reliable AI agents.
Bonus / Preferred Qualifications
  • Deep Telephony Infrastructure: Hands-on experience with bare-metal VoIP networks, custom SIP trunks, and RTP streaming. Familiarity with managing and configuring PBX systems like Asterisk or FreeSWITCH, and handling the network latency and jitter inherent to low-level telecom systems.
Ideal Technical Stack
  • Languages: Python, C++, Go (for high-performance backend components)
  • AI/ML: PyTorch, vLLM, HuggingFace, LangChain/LlamaIndex, specialized ASR/TTS frameworks
  • Telephony & Audio: WebRTC, standard telecom APIs (Twilio/Exotel), standard audio encoding (8kHz µ-law/A-law)
  • Backend & Infrastructure: Kubernetes, Docker, gRPC, Apache Kafka, Temporal, Redis, PostgreSQL
  • Observability: Prometheus, Grafana, OpenTelemetry (focusing on sub-millisecond tracing for audio/text pipelines)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Engineer – Agentic Systems & Voice AI
Lead AI Engineer – Agentic Systems & Voice AI

YAL • Hyderabad

On-site
INR 4,500,000 - 9,000,000
Lead AI Engineer – Agentic Systems & Voice AI
Lead AI Engineer – Agentic Systems & Voice AI

Keka Technologies Private Limited • Hyderabad

On-site
INR 3,500,000 - 7,000,000
Senior Ai Engineer Voice Ai Agentic Systems Danish Mullaji Gurugram
Senior Ai Engineer Voice Ai Agentic Systems Danish Mullaji Gurugram

Vibehackers • Gurugram District

On-site
INR 4,000,000 - 7,000,000
AI Engineer
AI Engineer

Technology services Company • Mumbai

On-site
INR 350,000 - 550,000
AI Agentic Engineer
AI Agentic Engineer

360 Degree Cloud • Dadri

On-site
INR 1,800,000 - 2,400,000
Applied Scientist-Agentic AI
Applied Scientist-Agentic AI

Magnet HR Tech Digital • Mumbai

On-site
INR 1,800,000 - 3,200,000
Lead AI Engineer
Lead AI Engineer

United States Digital Space LLC • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Staff AI Engineer — Agentic AI
Staff AI Engineer — Agentic AI

Gnani Innovations Private Limited. • India

On-site
INR 4,000,000 - 7,500,000
Team Lead - Conversational Agents(LLM & Agentic AI)
Team Lead - Conversational Agents(LLM & Agentic AI)

RingCentral • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Back End Developer
Back End Developer

Recrew AI • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Opportunity to build impactful voice AI infrastructure
Access to next-gen GPU clusters
Competitive compensation and benefits