Sr Site Reliability Engineer

SkanAI

Karnataka

On-site

INR 1,200,000 - 1,600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Skan AI is seeking an experienced Site Reliability Engineer to own the reliability, availability, and operational health of our cloud-hosted customer environments. You will bridge software and infrastructure, applying engineering discipline to automate toil, respond to incidents rapidly, and uphold SLA commitments for enterprise clients.

As part of the Customer Cloud Ops team, you’ll participate in on-call rotations, design agentic workflows, and collaborate with DevOps and product teams to

Qualifications

  • 5-8 years of experience as a SRE engineer.
  • Cloud platforms: AWS, Azure, or GCP — environment management, networking, IAM, and observability at scale.
  • Observability and monitoring: Prometheus, Grafana, Datadog, or equivalent — building dashboards, alerts, and SLO tracking.
  • Infrastructure as Code: Terraform, Ansible, or Pulumi — provisioning and configuration management.
  • Incident management: on-call discipline, structured MTTR mindset, post-mortem culture, and blameless review practices.
  • Scripting and automation: Python, Bash — automation of operational tasks and agentic workflow development.
  • Linux systems administration: process management, log analysis, performance tuning.
  • Container orchestration: Kubernetes and Docker — deployment management and debugging in production.
  • SRE fundamentals: SLI/SLO/SLA definition, error budget management, toil measurement and reduction.
  • Strong written documentation skills — clear, evidence-based runbooks and incident reports.

Responsibilities

  • Monitor platform health and performance across all cloud-hosted customer environments using observability tooling.
  • Respond to and own P1/P2 incidents — lead triage, diagnosis, and resolution; drive MTTR reduction through post-incident reviews.
  • Perform capacity and environment planning for customer cloud deployments — anticipate growth and prevent resource-related incidents.
  • Design and implement SRE automation to eliminate repetitive toil — agentic alerting, auto-remediation scripts, and automated runbook execution.
  • Manage change and release events that affect production customer environments — coordinate with DevOps and product teams to minimize risk.
  • Maintain and improve runbooks for all known failure patterns and operational procedures.
  • Contribute to HA/DR playbook validation.
  • Participate in on-call rotation and respond to alerts within defined SLA windows.
  • Track and report on SLO/SLA adherence — contribute to monthly operational reports and QBR data.
  • Identify and elevate environment risks proactively — before they become customer-facing incidents.
  • Collaborate with the Automation Engineering team to develop agentic workflows that automate triage, routing, and remediation.

Skills

SRE engineering
Cloud platforms AWS/Azure/GCP
Observability
Infrastructure as Code
Incident management
Scripting Python Bash
Linux administration
Kubernetes & Docker
SLI/SLO/SLA
Documentation

Tools

Terraform
Ansible
Pulumi
Prometheus
Grafana
Datadog
Kubernetes
Docker

Job description

Be at the Forefront of the Agentic AI Revolution

At Skan AI, you'll be part of the team pioneering the context engine for human and agentic execution, bringing context from enterprise operators, systems, and processes to power how the world's largest organizations execute their most complex, mission-critical work.


We're in hyper-growth mode at exactly the right moment in history. As enterprises race to adopt agentic AI, we're uniquely positioned to deliver the clear signal they desperately need: a platform that trains and grounds AI Agents in trillions of real execution signals, enabling reliable, compliant automation of their most complex processes.


Backed by Dell Technologies Capital and other leading investors, we're the only company that can bridge the gap between AI's promise and enterprise reality, making us perfectly positioned to define the agentic era for modern enterprises.


Our diverse, collaborative team of 250+ innovators is solving category-defining challenges at the intersection of AI, process intelligence, and enterprise work. Diverse perspectives fuel breakthrough thinking, cross-functional collaboration is the norm, and our work directly transforms how Fortune 500 companies operate. We are shaping the future of work itself.


The Site Reliability Engineer (SRE) owns the reliability, availability, and operational health of Skan's cloud-hosted customer environments. Sitting within the Customer Cloud Ops team, this role bridges software and infrastructure — applying engineering discipline to automate toil, respond to incidents rapidly, and uphold the SLA commitments that protect Skan's enterprise relationships.


The SRE is accountable for keeping production customer environments running at the performance and availability standards enterprise clients expect. This means being on-call, being proactive about operational risk, and continuously eliminating the manual work that gets in the way of reliable operations.


WHY THIS ROLE EXISTS

As Skan's cloud-hosted customer base grows, the complexity and volume of environment management, incident response, and reliability engineering grows with it. Without dedicated SRE capability, operational toil accumulates, incidents take longer to resolve, and SLA commitments are at risk — damaging customer trust and creating costly escalations.


The SRE function applies software engineering practices to infrastructure and operations — replacing manual, reactive processes with automation, runbooks, and agentic workflows that scale.


KEY RESPONSIBILITIES


  • Monitor platform health and performance across all cloud-hosted customer environments using observability tooling (Prometheus, Grafana, Datadog, or equivalent)

  • Respond to and own P1/P2 incidents — lead triage, diagnosis, and resolution; drive MTTR reduction through structured post-incident review

  • Perform ongoing capacity and environment planning for customer cloud deployments — anticipating growth and preventing resource-related incidents

  • Design and implement SRE automation to eliminate repetitive operational toil — agentic alerting, auto-remediation scripts, and automated runbook execution

  • Manage change and release events that affect production customer environments — coordinating with DevOps and product teams to minimize risk

  • Maintain and improve runbooks for all known failure patterns and operational procedures

  • Contribute to HA/DR playbook validation

  • Participate in on-call rotation and respond to alerts within defined SLA windows

  • Track and report on SLO/SLA adherence — contributing to monthly operational reports and QBR data

  • Identify and elevate environment risks proactively — before they become customer-facing incidents

  • Collaborate with the Automation Engineering team to develop agentic workflows that automate triage, routing, and remediation


KEY DELIVERABLES


  • Incident response records and post-mortems for all P1/P2 events — with root cause, remediation steps, and prevention actions documented

  • SLA/SLO dashboards: availability and performance reports for all cloud customer environments — updated continuously

  • Capacity and environment sizing plans per customer — reviewed quarterly or on significant usage change

  • Automated toil-reduction scripts and agentic remediation workflows — measurable reduction in manual operational hours

  • Runbooks for all known failure patterns — maintained and validated against real incidents

  • Change management records for all production environment events


KEY SKILLS & QUALIFICATIONS


  • 5-8 years of experience as a SRE engineer

  • Cloud platforms: AWS, Azure, or GCP — environment management, networking, IAM, and observability at scale

  • Observability and monitoring: Prometheus, Grafana, Datadog, or equivalent — building dashboards, alerts, and SLO tracking

  • Infrastructure as Code: Terraform, Ansible, or Pulumi — provisioning and configuration management

  • Incident management: on-call discipline, structured MTTR mindset, post-mortem culture, and blameless review practices

  • Scripting and automation: Python, Bash — automation of operational tasks and agentic workflow development

  • Linux systems administration: process management, log analysis, performance tuning

  • Container orchestration: Kubernetes and Docker — deployment management and debugging in production

  • SRE fundamentals: SLI/SLO/SLA definition, error budget management, toil measurement and reduction

  • Strong written documentation skills — clear, evidence-based runbooks and incident reports


Skan AI is an equal opportunity employer committed to building a diverse, inclusive, and respectful workplace around the world. We do not discriminate based on race, color, religion or belief, sex (including pregnancy, sexual orientation, gender identity, or gender expression), national origin, ancestry, age, disability, medical condition, genetic information, marital or family status, military or veteran status, or any other characteristic protected by applicable laws in the locations where we operate.


We welcome people from all backgrounds and provide reasonable accommodations throughout the hiring process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Skan AI • Bengaluru

Hybrid
INR 4,200,000 - 6,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer — Multi-Cloud Infrastructure
Site Reliability Engineer — Multi-Cloud Infrastructure

Skit.ai • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer - Level - 3
Site Reliability Engineer - Level - 3

CorroHealth • Dadri

On-site
INR 900,000 - 1,800,000
Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

Hybrid
INR 1,500,000 - 2,300,000
Associate Principal Site Reliability Engineer
Associate Principal Site Reliability Engineer

Saviynt • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Sr. Director, AI Solutions
Sr. Director, AI Solutions

Skanai • India

On-site
INR 5,500,000 - 7,500,000
Sr. Director, AI Solutions
Sr. Director, AI Solutions

SkanAI • Karnataka

On-site
INR 4,000,000 - 7,000,000
Sr. Director, AI Transformation & Enterprise Value
Sr. Director, AI Transformation & Enterprise Value

SkanAI • Karnataka

On-site
INR 19,375,000 - 24,541,000