Senior Site Reliability Engineer

iLink Digital

Chennai

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A technology services company is seeking an experienced SRE/DevOps Engineer with 6-10 years of experience in managing AWS and Azure environments. The role involves designing Kubernetes clusters, leading incident management, and collaborating across teams for operational improvements. Candidates should be well-versed in monitoring and observability tools and have a strong understanding of infrastructure as code practices. Excellent communication skills are essential for cross-team collaboration.

Qualifications

  • 6–10 years of experience in SRE, DevOps, or production support engineering.
  • Proven experience managing multi-cloud environments (AWS + Azure).
  • Familiarity with infrastructure as code (IaC) best practices.

Responsibilities

  • Design, deploy, and manage Kubernetes clusters for production workloads.
  • Lead incident management for production issues across environments.
  • Collaborate with development teams on operational and performance improvements.

Skills

SRE
DevOps
AWS
Azure
Kubernetes
PostgreSQL
Terraform
Ansible
Incident Management
Monitoring Tools (Prometheus, Grafana)

Tools

Prometheus
Grafana
Datadog
Splunk
ELK
Helm
ArgoCD

Job description

Requirements
  • 6–10 years of experience in SRE, DevOps, infrastructure and production support engineering roles.
  • Proven experience managing multi-cloud environments (AWS + Azure).
  • Demonstrated experience handling P1/P2 production incidents in cloud environments.
  • Familiarity with Prometheus, Grafana, Datadog, or Splunk.
  • Design, deploy, and manage Kubernetes clusters for production workloads at scale.
  • Architect and maintain PostgreSQL databases — performance tuning, HA setup, backup/restore strategies.
  • Build and manage cloud infrastructure on AWS and Azure using Terraform and Ansible.
  • Lead vulnerability management programs — identify, prioritize, and remediate security risks across the stack.
  • Define and enforce SLOs, SLIs, and error budgets; drive reliability improvements across services.
  • Implement IaC best practices, automate provisioning pipelines, and reduce manual toil.
  • Collaborate with development teams on capacity planning, disaster recovery, and incident post-mortems.
  • Build and maintain monitoring, alerting, and observability frameworks (Prometheus, Grafana, ELK, etc.).
  • Lead end-to-end incident management — detection, triage, escalation, resolution, and communication.
  • Serve as an on-call engineer; manage and respond to alerts and production incidents effectively.
  • Conduct blameless post-mortems and implement action items to prevent recurrence.
  • Monitor system health using dashboards and alerting tools; proactively identify degradation risks.
  • Collaborate with Dev, QA, and infrastructure teams to identify and reduce toil and failure points.
  • Support Kubernetes workloads and assist in troubleshooting cluster-level issues.
  • Work across AWS and Azure environments for incident containment and recovery.
  • Maintain and improve runbooks, playbooks, and incident response documentation.
  • Strong understanding of networking, security, and distributed systems.
  • Excellent communication skills for cross-team collaboration and post-mortem documentation.
  • Experience with Helm, ArgoCD, or GitOps workflows.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Senior DevSecOps/Site Reliability Engineer (AWS)
Senior DevSecOps/Site Reliability Engineer (AWS)

Stryker Group • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

Saika Technologies Inc. • Hyderabad, Bengaluru

Hybrid
INR 3,000,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

Recro • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Skillventory • Kamrup Metropolitan

On-site
INR 1,400,000 - 2,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Embarkgcc Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BuildxPartners • Bengaluru Urban

Hybrid
INR 2,400,000 - 4,200,000