Get more replies from employers
Send a job-specific resume in minutes.
Skit.ai seeks a Site Reliability Engineer to own the day-to-day reliability of a multi-cloud infrastructure spanning AWS, GCP, and Azure. You will manage compute, storage, networking, and identity with IaC, CI/CD, and Kubernetes, while driving automation and on-call incident response to keep services healthy.
The role emphasizes cost controls, backups, secrets management, and scalable self-service tooling. Strong Linux, networking, and automation skills are essential for success.
5+ years
Full-Time
Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/
Job Title: Site Reliability Engineer — Multi-Cloud Infrastructure
Type: Full-time
We run a voice AI platform for regulated enterprises in banking, telecom, and collections, spread across AWS, GCP, and Azure — for resilience, for cost, and because client data‑residency rules leave us no choice. That's a lot of surface area: compute, networking, storage, identity, clusters, pipelines, and supporting services, all needing to stay healthy across three providers.
This role owns the day-to-day reliability and operations of that estate. It's the generalist counterpart to our real-time-platform SRE: where they go deep on the latency‑critical call path, you go broad — keeping the whole infrastructure dependable, well‑automated, and cost‑sane, and sharing the on‑call load. If you like knowing how everything fits together and making the boring parts reliable and self‑serve, this is a good seat.
Representative — you'll help shape it. Multi‑cloud across AWS, GCP, and Azure; Kubernetes/containers; Terraform and GitHub Actions CI/CD; PostgreSQL; Grafana/Tempo for monitoring; Modal for ML deployment; LiveKit/SIP telephony on the platform side.
The infrastructure just works, across all three clouds, and when it doesn’t it's caught early and fixed cleanly. Engineers provision what they need without filing tickets. Cloud spend is understood, not surprising. And the on‑call rotation trends calmer because the estate is increasingly automated and self‑healing.
We’re an equal‑opportunity employer and evaluate every candidate on merit.