Senior Cloud Reliability Engineer

Mode Analytics

Thiruvananthapuram

On-site

INR 7,000,000 - 12,000,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mode Analytics is seeking a Senior Cloud Reliability Engineer to own availability, reliability, and security for a large-scale multi-cloud SaaS platform (AWS, GCP). You will operate Kubernetes-based control/data planes, implement AI-augmented automation, and drive incident response with blameless post-mortems.

You bring 6+ years in Enterprise SaaS Ops, Go/Python programming, and IaC experience with Terraform/Ansible, plus strong cloud security and observability skills.

Qualifications

  • 6+ years of Enterprise SaaS Ops experience.
  • Proficiency in Go and Python; IaC with Terraform and Ansible.
  • Expertise in Cloud Security and/or Cloud networking.
  • Experience with AI Ops tools and Agentic LLM.
  • Knowledge of cloud services, Kubernetes, cloud databases, Elasticsearch/Opensearch, and microservice architecture.
  • Experience implementing observability stacks (metrics, logs, tracing) and alerting in Cloud SaaS environments.
  • Strong debugging and problem-solving across network, systems, database, and application layers.
  • Certifications from AWS, Azure, GCP are a bonus.

Responsibilities

  • Operate a high-scale multi-cloud SaaS platform (AWS, GCP) ensuring reliability, performance, and uptime for production workloads.
  • Embed AI-driven workflows into SRE practices for anomaly detection, automated triage, runbook execution, and incident summarization to reduce MTTR.
  • Lead capacity planning and scaling with autoscaling strategies to support SaaS growth.
  • Architect and operate Kubernetes controller frameworks and enforce standards for cluster lifecycle and failover.
  • Own operations of cloud-native databases and data infrastructure, including performance tuning and backup/recovery.
  • Lead incident response and blameless post-mortems to permanent resolutions and prevention.
  • Promote automation-first culture with self-healing systems and IaC tooling.

Skills

Go
Python
SRE/Production Ops
AI Ops
Observability
Incident response
Troubleshooting
On-call leadership
Cloud security
Kubernetes

Education

B.Tech. degree in Computer Science or equivalent

Tools

Terraform
Ansible
Kubernetes
Helm
GitOps
AWS
GCP
PostgreSQL
DynamoDB
MySQL
Elasticsearch/OpenSearch

Job description

Senior Cloud Reliability Engineer

We are seeking a Senior Cloud Reliability Engineer with deep enterprise SaaS operations expertise to own the availability, reliability, security, and efficiency of our Multi-Cloud (AWS, GCP) production SaaS platform. The ideal candidate brings hands-on experience running highly available, large-scale Kubernetes-based control and data planes, a strong bias toward automation and AI-augmented operations, and a proven track record in production security, capacity management, and cloud-native data infrastructure.

Responsibilities
  • Operate a high-scale, multi-cloud (AWS, GCP) SaaS platform — ensuring reliability, performance, and uptime for business-critical production workloads.
  • Embed AI and Agentic workflows into SRE practice: leverage AI Ops platforms and LLM-powered autonomous agents for anomaly detection, automated triage, runbook execution, and incident summarization to reduce MTTR.
  • Drive capacity planning and scaling operations — proactively model growth, right-size infrastructure, and implement horizontal/vertical autoscaling strategies to support SaaS growth without reliability regression.
  • Architect and operate Kubernetes controller frameworks governing both control plane and data plane services; define and enforce operational standards for cluster lifecycle, workload scheduling, autoscaling, and failover.
  • Own operations of high-scale cloud-native databases and data infrastructure: PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/OpenSearch, ElastiCache (Redis/Memcached) on AWS and GCP — including performance tuning, backup/recovery, and incident response.
  • Lead incident response and blameless post-mortems for P0/P1 events; drive root cause analysis to permanent resolution and prevention — eliminating repeat incidents through systemic fixes, not workarounds.
  • Define and enforce a culture of automation-first: identify and eliminate toil through self-healing systems, automated remediation pipelines, and infrastructure-as-code (Terraform, Helm, GitOps).
  • Participate in on-call rotations for critical cloud infrastructure; serve as a senior escalation point and incident commander during high-severity events.
  • Achieve quantifiable SaaS operational Excellence measured by related SLI/SLO/SLA
Required skills/qualifications
  • B.Tech. degree in Computer Science or equivalent.
  • At least 6+ years of Enterprise SaaS Ops experience
  • Strong proficiency in programming, particularly with Go and Python, and experience with Infrastructure as Code (IaC) tools like Terraform and Ansible.
  • Expertise in Cloud Security and/or Cloud networking
  • Experience with AI Ops tools, Agentic LLM.
  • Experience/ Knowledge in Cloud Services, Kubernetes, Cloud Databases like Postgres/RDS/MySQL/DynamoDB, Elastic, Kafka, and Microservice architecture is a bonus.
  • Experience in implementing and operating enterprise-grade observability ( metrics, logs, tracing), alerting stack in a Cloud SaaS environment
  • Strong debugging and problem-solving skills (network, systems, database, and application).
  • Advanced professional certifications from Cloud Providers ( AWS, Azure, GCP) in domains like K8s, Solution architecture, networking, and databases are a bonus.
  • Full Stack Architecture/Development Experience is a bonus.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Reliability Engineer
Senior Cloud Reliability Engineer

ThoughtSpot • Thiruvananthapuram

On-site
INR 2,800,000 - 4,600,000
Senior SRE Engineer
Senior SRE Engineer

TymblHub • Chennai District

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer_ AWS certified
Site Reliability Engineer_ AWS certified

PwC Acceleration Center India • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Cloud Engineer
Senior Cloud Engineer

LE300 Optiva (India) Technologies Pvt. Ltd. • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Cloud & SecOps Engineer
Cloud & SecOps Engineer

NowFloats • Hyderabad

On-site
INR 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BuildxPartners • Bengaluru Urban

Hybrid
INR 2,400,000 - 4,200,000
Senior DevOps Engineer
Senior DevOps Engineer

Unico Connect LLP. • Mumbai

On-site
INR 1,500,000 - 2,500,000
Cloud Architect and SRE Lead
Cloud Architect and SRE Lead

Hector And Streak Consulting • Mumbai Suburban

Hybrid
INR 4,000,000 - 7,000,000