Site Reliability Engineer — Multi-Cloud Infrastructure

Skit.ai

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Skit.ai seeks a Site Reliability Engineer to own the day-to-day reliability of a multi-cloud infrastructure spanning AWS, GCP, and Azure. You will manage compute, storage, networking, and identity with IaC, CI/CD, and Kubernetes, while driving automation and on-call incident response to keep services healthy.

The role emphasizes cost controls, backups, secrets management, and scalable self-service tooling. Strong Linux, networking, and automation skills are essential for success.

Qualifications

  • Must have SRE/DevOps experience in production systems with on-call responsibility.
  • Hands-on across at least two cloud providers (AWS,GCP,Azure).
  • Production Kubernetes and container experience.
  • Infrastructure-as-code using Terraform or similar tools.
  • Strong CI/CD experience and clear deployment automation.
  • Experience with monitoring/alerting: Prometheus/Grafana.
  • Automation scripting in Python/Go/Bash.
  • Solid Linux systems and networking fundamentals.

Responsibilities

  • Own multi-cloud operations to keep compute, networking, storage, and identity healthy across AWS, GCP, and Azure.
  • Manage IaC, CI/CD pipelines, and ensure reproducible, reviewed changes.
  • Maintain Kubernetes clusters and supporting services (databases, queues, caches).
  • Set up and improve monitoring/alerting, participate in on-call, and lead incident response.
  • Drive cost hygiene, patching, backups, secrets management, and access controls.
  • Automate repetitive tasks and build self-service tooling to scale the estate.

Skills

SRE/DevOps experience
Cloud platforms (AWS,GCP,Azure)
Kubernetes
Containers
Terraform / IaC
CI/CD pipelines
Prometheus
Grafana
Python
Go
Bash
Linux administration
Networking fundamentals
PostgreSQL
Security best practices

Job description

5+ years

Full-Time

About the Role

Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/

Job Title: Site Reliability Engineer — Multi-Cloud Infrastructure

Type: Full-time

Why this role exists

We run a voice AI platform for regulated enterprises in banking, telecom, and collections, spread across AWS, GCP, and Azure — for resilience, for cost, and because client data‑residency rules leave us no choice. That's a lot of surface area: compute, networking, storage, identity, clusters, pipelines, and supporting services, all needing to stay healthy across three providers.

This role owns the day-to-day reliability and operations of that estate. It's the generalist counterpart to our real-time-platform SRE: where they go deep on the latency‑critical call path, you go broad — keeping the whole infrastructure dependable, well‑automated, and cost‑sane, and sharing the on‑call load. If you like knowing how everything fits together and making the boring parts reliable and self‑serve, this is a good seat.

What you'll own
  • Multi‑cloud operations. Provision, operate, and keep healthy compute, networking, storage, and identity across AWS, GCP, and Azure — with sensible consistency instead of three snowflakes.
  • Infrastructure as code. Manage the estate through Terraform (or equivalent) and version control — reproducible environments, reviewed changes, no undocumented hand‑tweaks.
  • CI/CD and delivery. Keep build and deploy pipelines fast and reliable so engineers ship safely and often.
  • Clusters and workloads. Run Kubernetes/container platforms and the supporting services (databases, queues, caches, internal tooling) that everything depends on.
  • Monitoring and on‑call. Maintain monitoring and alerting for infrastructure health, take a turn in the rotation, and respond to and mitigate incidents with clear communication and blameless follow‑up.
  • Cost and hygiene. Keep an eye on cloud spend, rightsizing, and waste; own the unglamorous but essential hygiene — patching, backups, secrets, and access.
  • Automation and toil reduction. Replace manual, repetitive operations with automation and self‑service so the team scales without headcount scaling with it.
What the first year looks like
  • First 90 days. Learn the estate across all three clouds. Take a turn on call. Close the most obvious gaps in monitoring, backups, and access hygiene.
  • By 6 months. More of the estate under consistent infrastructure‑as‑code. Reliable, reviewed CI/CD. A clearer, quieter alerting setup and documented runbooks for the common incidents.
  • By 12 months. Measurably less manual toil through automation and self‑service. Sensible cost controls in place. Provisioning and environment setup that's repeatable rather than tribal knowledge.
What we’re looking for
Must‑have
  • A few years in SRE, DevOps, or infrastructure operations for production systems, including on‑call.
  • Hands‑on experience across at least two of AWS, GCP, and Azure (all three is a strong plus).
  • Kubernetes and containers in production.
  • Infrastructure‑as‑code (Terraform or similar) and CI/CD pipelines.
  • Monitoring and alerting practice (e.g. Prometheus/Grafana) and structured incident handling.
  • A scripting/programming language for automation (Python, Go, or Bash beyond one‑liners).
  • Solid Linux systems and networking fundamentals.
Nice‑to‑have
  • All three clouds run in production, and comfort designing for consistency across them.
  • Secrets management, security hardening, and compliance/data‑residency contexts.
  • PostgreSQL and other stateful‑service operations at scale.
  • Some exposure to real‑time or voice infrastructure — enough to back up the platform SRE on call.
Our Stack

Representative — you'll help shape it. Multi‑cloud across AWS, GCP, and Azure; Kubernetes/containers; Terraform and GitHub Actions CI/CD; PostgreSQL; Grafana/Tempo for monitoring; Modal for ML deployment; LiveKit/SIP telephony on the platform side.

How you’ll know you’re succeeding

The infrastructure just works, across all three clouds, and when it doesn’t it's caught early and fixed cleanly. Engineers provision what they need without filing tickets. Cloud spend is understood, not surprising. And the on‑call rotation trends calmer because the estate is increasingly automated and self‑healing.

We’re an equal‑opportunity employer and evaluate every candidate on merit.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer Voice AI Platform
Senior Site Reliability Engineer Voice AI Platform

Skit.ai • Bengaluru

On-site
INR 4,200,000 - 6,200,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Sr Site Reliability Engineer
Sr Site Reliability Engineer

SkanAI • Karnataka

On-site
INR 1,200,000 - 1,600,000
Director Cloud & Infrastructure Architect (Multi-Cloud | Datacenter | SRE)
Director Cloud & Infrastructure Architect (Multi-Cloud | Datacenter | SRE)

Mancer Consulting Services • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

WorkSpan • Bengaluru

On-site
INR 2,200,000 - 3,600,000
Platform Engineer
Platform Engineer

United States Digital Space LLC • Maharashtra

On-site
INR 1,500,000 - 2,800,000
Site Reliability Engineer
Site Reliability Engineer

DeepIQ • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,200,000
Significant equity in a venture-backed company
Opportunity to work with modern tech stack
Site Reliability Engineer(SRE)
Site Reliability Engineer(SRE)

Techdome • Hyderabad

On-site
INR 150,000 - 210,000