Senior Backend Engineer - Speech Platform

Zoho

Dadri

On-site

INR 1,400,000 - 2,400,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Zoho is seeking an experienced backend engineer to own our speech team’s systems, from ingestion to real-time serving and integration with our Java-based platform. You will build data pipelines, streaming inference services, and training infrastructure, while optimizing latency and compute cost.

You’ll collaborate with ML engineers to deploy models at scale, handle on-premise deployments for restricted data environments, and ensure robust telemetry and performance profiling across the stack.

Qualifications

  • Must have 3–4 years of building and operating production backend systems.
  • Strong Java—you've owned services in production, not just contributed to them.
  • Working Python—enough to build data pipelines and integrate with ML tooling.
  • Real-time or low-latency systems experience: streaming APIs, WebSocket or gRPC streaming, concurrency, backpressure.
  • Data pipelines at scale (Airflow, Dagster, Spark or equivalent).
  • Docker and Kubernetes in production.
  • Cloud infrastructure (AWS/GCP/Azure).
  • Comfortable debugging performance: profiling, latency percentiles, throughput under load.

Responsibilities

  • Build the audio data pipeline: mine our call archive, transcode, resample, segment, deduplicate and quality-filter at scale.
  • Build and operate real-time inference services with hard latency targets, including streaming, cancellation and mid-utterance interruption.
  • Integrate speech services into our existing Java-based platform and telephony infrastructure.
  • Instrument the full latency budget end to end and find where the milliseconds go.
  • Stand up training infrastructure — GPU scheduling, checkpointing, experiment tracking, reproducibility.
  • Own compute cost and concurrency economics: how many simultaneous calls per GPU, and how to improve it.
  • Support on-premise deployment for clients who can't send data outside their network.

Skills

Java
Python
Real-time systems
Data pipelines
Airflow
Dagster
Spark
Docker
Kubernetes
AWS
GCP
Azure

Tools

Airflow
Dagster
Spark
Docker
Kubernetes

Job description

You'll be the engineering owner on the speech team, working alongside two ML engineers and building everything around the models — ingestion, training infrastructure, real-time serving, and the integration into our existing platform.

This is a backend engineering role. You don't need to train models. You need to build the systems that make trained models useful in production.

Responsibilities:

  • Build the audio data pipeline: mine our call archive, transcode, resample, segment, deduplicate and quality-filter at scale
  • Build and operate real-time inference services with hard latency targets, including streaming, cancellation and mid-utterance interruption
  • Integrate speech services into our existing Java-based platform and telephony infrastructure
  • Instrument the full latency budget end to end and find where the milliseconds go
  • Stand up training infrastructure — GPU scheduling, checkpointing, experiment tracking, reproducibility
  • Own compute cost and concurrency economics: how many simultaneous calls per GPU, and how to improve it
  • Support on-premise deployment for clients who can't send data outside their network
Requirements

Must haves:

  • 3–4 years building and operating production backend systems
  • Strong Java — you've owned services in production, not just contributed to them
  • Working Python — enough to build data pipelines and integrate with ML tooling
  • Real-time or low-latency systems experience: streaming APIs, WebSocket or gRPC streaming, concurrency, backpressure
  • Data pipelines at scale (Airflow, Dagster, Spark or equivalent)
  • Docker and Kubernetes in production
  • Cloud infrastructure (AWS/GCP/Azure)
  • Comfortable debugging performance: profiling, latency percentiles, throughput under load
Nice to have:
  • Serving ML models in production (Triton, vLLM, TorchServe)
  • Telephony — SIP, Asterisk/FreeSWITCH, media servers, narrowband codecs
  • MLOps tooling: MLflow, Weights & Biases, DVC
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Backend Engineer - Speech Platform
Senior Backend Engineer - Speech Platform

Gohyred • Uttar Pradesh

On-site
INR 1,200,000 - 2,000,000
ML Engineer - Speech
ML Engineer - Speech

Zoho • Dadri

On-site
INR 1,200,000 - 2,400,000
ML Research Engineer Speech
ML Research Engineer Speech

Blue Machines AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
ML Research Engineer, Speech
ML Research Engineer, Speech

Blue Machines AI • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization
Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization

OutcomesAI • Bengaluru

On-site
INR 1,500,000 - 2,000,000
ML Engineer - Speech
ML Engineer - Speech

Gohyred • Uttar Pradesh

On-site
INR 1,500,000 - 2,200,000
Senior ML Research Scientist, Speech
Senior ML Research Scientist, Speech

Blue Machines AI • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Backend Engineer
Backend Engineer

VoisX • Bengaluru

On-site
INR 1,200,000 - 1,800,000
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
Senior Audio ML / Research Engineer
Senior Audio ML / Research Engineer

Mowka • Bengaluru

On-site
INR 6,000,000 - 8,000,000