Senior Site Reliability Engineer Voice AI Platform

Skit.ai

Bengaluru

On-site

INR 4,200,000 - 6,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Skit.ai is looking for a Senior Site Reliability Engineer to own the reliability and performance of our Voice AI Platform, handling real-time media, sub-second budgets, and six-figure concurrency across multiple clouds. You’ll own SLOs, observability, capacity planning, and incident response for regulated enterprise workloads in banking, telecom, and collections.

You'll collaborate with an AI Observability team and drive automation through infrastructure-as-code, CI/CD, and self-healing systems.

Qualifications

  • Several years running high-availability, high-throughput production systems.
  • Depth in at least one major cloud (AWS, GCP, or Azure) and with containers/Kubernetes.
  • Strong observability practice — metrics, distributed tracing, and logging.
  • Fluency with SLIs/SLOs/error budgets and structured incident management.
  • Infrastructure-as-code (Terraform or similar) and CI/CD.
  • A programming language for automation and tooling (Python, Go, or similar).
  • Solid Linux systems and networking fundamentals; capacity planning and performance tuning.

Responsibilities

  • Define and defend SLOs and error budgets for availability and latency across the call path.
  • Own the metrics, tracing, and logging stack so failures surface fast.
  • Maintain the real-time media path (SIP signaling, RTP) at scale.
  • Plan for peak capacity and tune autoscaling to avoid dropped calls or idle costs.
  • Run a calm, structured on-call with blameless postmortems.
  • Design for multi-cloud resilience across AWS, GCP, and Azure.
  • Automate toil with infrastructure-as-code and self-healing systems.

Skills

High availability
High throughput
AWS
GCP
Kubernetes
Observability
SLIs/SLOs
Infrastructure as code
Python/Go
Linux networking
On-call ownership

Tools

Prometheus
Grafana
OpenTelemetry
Tempo
Jaeger
Terraform
GitHub Actions

Job description

Senior Site Reliability Engineer — Voice AI Platform

5+ Years

Full-Time

About the Role

Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/

Job Title: Senior Site Reliability Engineer — Voice AI Platform

Type: Full-time

Why this role exists:

We run a voice AI platform that places and answers up to ~1 million calls per hour for regulated enterprises in banking, telecom, and collections. Unlike most SaaS, our workload is real-time and conversational: every call is a live media session where an extra few hundred milliseconds anywhere in the ASR → LLM → TTS loop is the difference between a natural exchange and a caller hanging up. Traffic is also bursty — outbound campaigns spin up huge concurrency inside narrow calling windows — and it runs across multiple clouds for resilience and data residency.

We are hiring a Senior SRE to own the reliability and performance of that system: the SLOs, the observability that makes problems visible, the capacity that absorbs campaign spikes, and the incident response that keeps regulated clients online. This is a systems-reliability role — latency, uptime, saturation, and the health of the telephony and serving path. (Model quality and evaluation live with a separate AI Observability role; you\'ll partner with them, not own their signals.)

If you want reliability problems that are genuinely hard — real-time media, sub-second budgets, six-figure concurrency, multi-cloud failover — this is that.

What you\'ll own:

  • SLOs and error budgets. Define and defend service-level objectives for availability and latency across the call path, and use error budgets to steer the balance between shipping and stability.
  • Observability. Own the metrics, tracing, and logging stack so failures surface fast and root cause is minutes not hours — distributed traces across signaling, ASR, LLM, TTS, and infra, with dashboards and alerting that page on real problems and stay quiet otherwise.
  • The real-time media path. Keep SIP signaling and RTP media healthy at scale — concurrency, jitter, packet loss, session setup — and the reliability of the components that carry them.
  • Capacity and autoscaling. Plan for peak (campaign windows that push toward the platform\'s concurrency ceiling), pre-warm capacity ahead of demand, and tune autoscaling so we neither drop calls nor burn money idling.
  • Incident response. Run a calm, structured on-call: triage, mitigation, clear comms to stakeholders on regulated accounts, and blameless postmortems that actually change the system.
  • Multi-cloud resilience. Design for failure across AWS, GCP, and Azure — redundancy, failover, disaster recovery, and the data-residency constraints that come with Indian banking and telecom clients.
  • Automation and toil reduction. Turn manual operations into infrastructure-as-code and self-healing systems. If you did it twice by hand, the third time is a script.

What the first your looks like:

  • First 90 days. Learn the call path end to end. Establish baseline SLIs for availability and latency, close the biggest gaps in alerting, and take a full turn in the on-call rotation.
  • By 6 months. Published SLOs with error budgets for the core services. A tracing/dashboards setup that makes the ASR→LLM→TTS latency budget visible per call. A repeatable pre-warm-and-scale playbook for campaign peaks.
  • By 12 months. Demonstrable reduction in incident frequency and time-to-mitigate. Tested multi-cloud failover for a critical path. On-call toil measurably down through automation.

What we\'re looking for

Must-have

  • Several years running high-availability, high-throughput production systems, including real on-call ownership.
  • Depth in at least one major cloud (AWS, GCP, or Azure) and with containers/Kubernetes.
  • Strong observability practice — metrics, distributed tracing, and logging (e.g. Prometheus/Grafana, OpenTelemetry, Tempo/Jaeger).
  • Fluency with SLIs/SLOs/error budgets and structured incident management.
  • Infrastructure-as-code (Terraform or similar) and CI/CD.
  • A programming language for automation and tooling (Python, Go, or similar) — beyond shell scripting.
  • Solid Linux systems and networking fundamentals; capacity planning and performance tuning.

Nice-to-have

  • Real-time media or VoIP experience — SIP/RTP, media servers, LiveKit, SBC/Kamailio.
  • Reliability of GPU/ML serving infrastructure.
  • Regulated-industry operations — uptime SLAs, DR, data residency.
  • Load testing and chaos engineering at scale.
  • PostgreSQL operations at scale.

Our stack:

Representative — you'll help shape it. Multi-cloud across AWS, GCP, and Azure; LiveKit/SIP for telephony; self-hosted and managed ASR (e.g. NVIDIA Parakeet / NeMo), LLMs, and TTS; Modal for ML deployment and pre-warming; PostgreSQL; Grafana/Tempo for metrics and traces; infrastructure-as-code and GitHub Actions CI/CD.

How you'll know you're succeeding:

Calls connect and stay fast even during the busiest campaign windows. Alerts mean something, and the ones that page you are worth waking up for. When something breaks, it's found and mitigated quickly and it doesn\'t break the same way twice. And the on-call rotation gets calmer over time, not busier, because the system increasingly heals itself.

We\'re an equal-opportunity employer and evaluate every candidate on merit. [Add benefits, compensation band, and application instructions before posting.]

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Multi-Cloud Infrastructure
Site Reliability Engineer — Multi-Cloud Infrastructure

Skit.ai • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Backend Engineer
Senior Backend Engineer

Skit.ai • Bengaluru

On-site
INR 12,000 - 18,000
Ownership Culture
Cutting-Edge AI Stack
Real Scale, Day One
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MyOperator • Dadri

On-site
INR 1,500,000 - 2,300,000
Senior backend Engineer
Senior backend Engineer

Infer • Karnataka

On-site
INR 1,000,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MyOperator • India

On-site
INR 1,200,000 - 1,800,000
Applied AI Engineer
Applied AI Engineer

Infer • Karnataka

On-site
INR 1,500,000 - 2,500,000
VP of Engineering
VP of Engineering

Gnani Innovations Private Limited. • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Lead/ Senior DevOps Engineer
Lead/ Senior DevOps Engineer

Skit.ai • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Staff Engineer, Product Infrastructure
Staff Engineer, Product Infrastructure

Sarvam AI • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Vahan.ai • Bengaluru

On-site
INR 4,000,000 - 7,000,000