Voice Engineer

OneByZero

Bengaluru

Hybrid

INR 2,800,000 - 6,000,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OneByZero in Bengaluru is building an AWS-native, real-time voice and AI platform for enterprises. You will own the ultra-low-latency audio architecture from design to production, bridging LLM orchestration with CTI and legacy telephony.

The role demands deep experience with WebRTC/WebSockets/gRPC, high-concurrency systems, and cloud-native deployments on AWS. You will ship a production-grade, multi-tenant voice pipeline and collaborate across teams to reduce latency and enable barrier-free

Qualifications

  • Experience building low-latency, real-time media servers at scale using open-source stacks.
  • Proficient in systems programming (Go, Rust, C++, Python) for concurrency and network I/O.
  • Experience with enterprise telephony: SIP, RTP, CTI protocols.
  • Knowledge of AWS networking and cloud-native deployments.
  • 5+ years in backend/infra/telephony engineering owning end-to-end systems.

Responsibilities

  • Architect the core audio pipeline with WebRTC, WebSockets, and gRPC for high concurrency.
  • Build CTI connectors to interface with legacy CCaaS platforms.
  • Maintain stateful streaming sessions on AWS with thousands of WebSocket/WebRTC connections.
  • Migrate the voice pipeline to production-ready, multi-tenant deployments using CI/CD and Terraform.
  • Optimize latency of audio transport, STT, and TTS handoffs to sub-500ms.
  • Develop Voice Activity Detection and endpointing to handle noise and barge-in.
  • Collaborate with product engineers to feed streams into LLM state management.

Skills

Low-latency media
Real-time streaming
Go
Rust
C++
Python
AWS networking
CTI integration
SIP trunking
WebRTC
WebSockets
gRPC
LiveKit
Kubernetes

Tools

LiveKit
Janus
Mediasoup
GStreamer
AWS
Kubernetes
Redis
Terraform
CI/CD

Job description

Core Product - Voice & Real-Time Infrastructure
About OneByZero

OneByZero builds AI-native platforms that help enterprises in regulated industries — banking,

telecommunications, and retail — move beyond AI experimentation into production-ready, governed AI workforces.

Our AWS-native platform deploys “governed AI coworkers”: autonomous agents that handle real business processes while maintaining full auditability, data sovereignty, and human oversight, typically moving from pilot to production in weeks rather than quarters.

We are not building a simple wrapper around a model — we are building a commercial platform that negotiates, handles complex workflows, and converses naturally with humans at enterprise scale.

Location: [Bangalore (India) / Hybrid] | Employment Type: Full-time | Team: Core Product Engineering
The Mission

At our core, we are solving for seamless human-machine interaction — using real-time voice as the ultimate interface. We are building an open-source-driven, AWS-native AI platform designed to deploy as an agentic digital coworker for enterprises. Your mission is to own the zero-to-one build-out of the ultra-low-latency audio architecture that makes this interaction feel instantaneous, bridging the gap between modern LLM orchestration and legacy enterprise telephony.

What You Will Do
  • Architect the core pipeline. Design and maintain a highly concurrent, bi-directional audio streaming infrastructure using WebRTC, WebSockets, and gRPC — handling the messy realities of network traversal (STUN/TURN, ICE candidate negotiation), packet loss, and transcoding relevant codecs (Opus, G.711 μ- law/A-law).
  • Bridge the legacy gap. Build the signaling and CTI (Computer Telephony Integration) connectors required to interface seamlessly with legacy on-prem and cloud CCaaS platforms (Avaya AES, Genesys, Cisco), managing call control signals to agent desktops alongside real-time RTP media extraction.
  • Manage state and concurrency at scale. Architect a distributed, cloud-native system on AWS (EKS, ElastiCache/Redis) capable of maintaining stateful streaming sessions and handling thousands of concurrent WebSocket/WebRTC connections without race conditions or dropped packets.
  • Drive the product rollout. Transition the voice pipeline from its current build into a hardened, production- ready system capable of supporting high-volume, multi-tenant enterprise deployments using modern CI/CD and Infrastructure as Code (Terraform).
  • Kill latency. Optimize every millisecond of the audio transport layer, STT, and TTS handoffs to keep system latency well under 500ms at scale.
  • Solve the “human” problems. Build robust Voice Activity Detection (VAD) and endpointing to handle natural pauses, background noise, and real-time human interruptions (barge-in).
  • Integrate with the “brain.” Work closely with our product engineers to ensure your audio streams and CTI signals feed cleanly into our LLM state management and tool-calling infrastructure.
What You Must Have
  • Deep, hands-on experience building low-latency, real-time media servers or audio pipelines at scale, heavily utilizing open-source frameworks (e.g., LiveKit, Janus, Mediasoup, or GStreamer).
  • Strong systems programming proficiency (Go, Rust, C++, or Python) for handling heavy concurrency and network I/O.
  • Knowledge of AWS networking and compute primitives, with a track record of deploying stateful, scalable streaming architectures natively in the cloud.
  • Proven experience wrestling with enterprise telephony: SIP trunking, RTP streams, and integrating with CTI protocols.
  • Expertise in managing ICE connection states, stream buffering, jitter buffers, and signal processing basics.
  • 5+ years of experience in backend, infrastructure, or media/telephony engineering, with a portion of that time spent owning a system end-to-end from design through production.
Nice to Have
  • Prior experience integrating speech-to-text (STT) and text-to-speech (TTS) providers into a real-time conversational pipeline.
  • Familiarity with LLM orchestration frameworks and function/tool-calling patterns.
  • Experience operating Kubernetes (EKS) in production, including autoscaling stateful workloads and zero- downtime deployments.
  • Exposure to observability tooling for real-time systems (distributed tracing, RTP/media quality metrics, latency dashboards).
  • Background in contact center technology, IVR systems, or telecom carrier integrations.
  • Experience contributing to or maintaining open-source real-time communication projects.
What Success Looks Like
  • A production-hardened voice pipeline running at sub-500ms end-to-end latency, supporting thousands of concurrent sessions across multiple enterprise tenants.
  • CTI connectors live with at least one major CCaaS platform (Avaya, Genesys, or Cisco), enabling seamless agent-desktop handoff.
  • A distributed session-state architecture on AWS that survives node failures and scales horizontally without dropped calls.
  • Natural, low-friction conversations — accurate barge-in handling and endpointing that feels human, not robotic.
Why Join OneByZero
  • Own a zero-to-one build: this is a foundational, high-ownership role shaping the core real-time infrastructure of the product.
  • Work at the intersection of telecom, distributed systems, and applied AI — solving problems most engineers never get exposed to.
  • Ship into real enterprise deployments in regulated industries, with direct visibility into production impact.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Voice Pipeline Engineer
Senior Voice Pipeline Engineer

OnebyZero . • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Back End Developer
Back End Developer

Recrew AI • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Opportunity to build impactful voice AI infrastructure
Access to next-gen GPU clusters
Competitive compensation and benefits
Lead AI Engineer
Lead AI Engineer

LeadSquared • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Back End Developer
Back End Developer

Super Humans • Bengaluru

On-site
INR 1,800,000 - 2,600,000
Senior Software Engineer, Voice AI
Senior Software Engineer, Voice AI

OmniDimension • India

On-site
INR 3,000,000 - 6,000,000
Full Stack Engineer (Voice AI)
Full Stack Engineer (Voice AI)

JoyzAI • Dadri

On-site
INR 1,800,000 - 2,600,000
Top-of-the-market pay
Work from office in Noida
Lead Artificial Intelligence Engineer
Lead Artificial Intelligence Engineer

Vahan.ai • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Infrastructure Engineer
Infrastructure Engineer

SourcingXPress • Surat

On-site
INR 700,000 - 1,000,000
WebRTC basics
Streaming media tooling
Always-on systems
Senior Voice AI Engineer — Convogent Delivery
Senior Voice AI Engineer — Convogent Delivery

Aivar Innovations • Coimbatore District

On-site
INR 1,500,000 - 2,200,000
Principal/ Staff Software Engineer, Voice AI
Principal/ Staff Software Engineer, Voice AI

OmniDimension • India

On-site
INR 4,000,000 - 7,000,000