Get more replies from employers
Send a job-specific resume in minutes.
Skit.ai is looking for a Senior Site Reliability Engineer to own the reliability and performance of our Voice AI Platform, handling real-time media, sub-second budgets, and six-figure concurrency across multiple clouds. You’ll own SLOs, observability, capacity planning, and incident response for regulated enterprise workloads in banking, telecom, and collections.
You'll collaborate with an AI Observability team and drive automation through infrastructure-as-code, CI/CD, and self-healing systems.
5+ Years
Full-Time
About the Role
Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/
Job Title: Senior Site Reliability Engineer — Voice AI Platform
Type: Full-time
Why this role exists:
We run a voice AI platform that places and answers up to ~1 million calls per hour for regulated enterprises in banking, telecom, and collections. Unlike most SaaS, our workload is real-time and conversational: every call is a live media session where an extra few hundred milliseconds anywhere in the ASR → LLM → TTS loop is the difference between a natural exchange and a caller hanging up. Traffic is also bursty — outbound campaigns spin up huge concurrency inside narrow calling windows — and it runs across multiple clouds for resilience and data residency.
We are hiring a Senior SRE to own the reliability and performance of that system: the SLOs, the observability that makes problems visible, the capacity that absorbs campaign spikes, and the incident response that keeps regulated clients online. This is a systems-reliability role — latency, uptime, saturation, and the health of the telephony and serving path. (Model quality and evaluation live with a separate AI Observability role; you\'ll partner with them, not own their signals.)
If you want reliability problems that are genuinely hard — real-time media, sub-second budgets, six-figure concurrency, multi-cloud failover — this is that.
What you\'ll own:
What the first your looks like:
What we\'re looking for
Must-have
Nice-to-have
Our stack:
Representative — you'll help shape it. Multi-cloud across AWS, GCP, and Azure; LiveKit/SIP for telephony; self-hosted and managed ASR (e.g. NVIDIA Parakeet / NeMo), LLMs, and TTS; Modal for ML deployment and pre-warming; PostgreSQL; Grafana/Tempo for metrics and traces; infrastructure-as-code and GitHub Actions CI/CD.
How you'll know you're succeeding:
Calls connect and stay fast even during the busiest campaign windows. Alerts mean something, and the ones that page you are worth waking up for. When something breaks, it's found and mitigated quickly and it doesn\'t break the same way twice. And the on-call rotation gets calmer over time, not busier, because the system increasingly heals itself.
We\'re an equal-opportunity employer and evaluate every candidate on merit. [Add benefits, compensation band, and application instructions before posting.]