Senior Site Reliability Engineer

iLink Digital

Chennai District

On-site

INR 1,800,000 - 3,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

iLink Digital is seeking an experienced SRE/DevOps engineer to join our Chennai-based delivery team. You will design and operate scalable cloud infrastructure, manage multi-cloud environments (AWS + Azure), implement CI/CD pipelines, observe systems with Prometheus, Grafana, and ELK stacks, and drive reliability improvements across services.

Ideal candidates have hands-on incident management experience, strong automation skills with Terraform and Ansible, and a passion for building resilient

Skills

SRE
DevOps
Incident management
Multi-cloud experience

Tools

Kubernetes
Terraform
Ansible
Prometheus
Grafana
Datadog
Splunk
ELK
Helm
ArgoCD
GitOps

Job description

iLink Digital is aGlobal Software Solution Provider and SystemsIntegrator, delivers next-generation technologysolutions to help clients solve complex businesschallenges, improve organizational effectiveness,increase business productivity, realize sustainableenterprise value and transform your businessinside-out. iLink integrates software systems and developscustom applications, components, and frameworks on the latestplatforms for IT departments, commercial accounts, applicationservices providers (ASP) and independent software vendors(ISV). iLink solutions are used in a broad range of industriesand functions, including healthcare, telecom, government, oiland gas, education, and life sciences. iLink’s expertiseincludes Cloud Computing & Application Modernization, DataManagement & Analytics, Enterprise Mobility, Portal,collaboration & Social Employee Engagement, EmbeddedSystems and User Experience designetc.

What makesiLink's offerings unique is the fact that we usepre-created frameworks, designed to accelerate softwaredevelopment and implementation of business processes for ourclients. iLink has over 60 frameworks (solution accelerators),both industry-specific and horizontal, that can be easilycustomized and enhanced to meet your current businesschallenges.

Requirements
  • 6–10 years of experience in SRE, DevOps,infrastructure and production support engineering roles.
  • Proven experience managing multi-cloudenvironments (AWS + Azure).
  • Demonstrated experience handling P1/P2 productionincidents in cloud environments.
  • Familiarity with Prometheus, Grafana, Datadog, orSplunk
  • Design, deploy, and manage Kubernetes clustersfor production workloads at scale.
  • Architect and maintain PostgreSQL databases —performance tuning, HA setup, backup/restore strategies.
  • Build and manage cloud infrastructure on AWS andAzure using Terraform and Ansible.
  • Lead vulnerability management programs —identify, prioritize, and remediate security risks across the stack.
  • Define and enforce SLOs, SLIs, and error budgets;drive reliability improvements across services.
  • Implement IaC best practices, automateprovisioning pipelines, and reduce manual toil.
  • Collaborate with development teams on capacityplanning, disaster recovery, and incident post-mortems.
  • Build and maintain monitoring, alerting, andobservability frameworks (Prometheus, Grafana, ELK, etc.).
  • Lead end-to-end incident management — detection,triage, escalation, resolution, and communication.
  • Serve as an on-call engineer; manage and respondto alerts and production incidents effectively.
  • Conduct blameless post-mortems and implementaction items to prevent recurrence.
  • Monitor system health using dashboards andalerting tools; proactively identify degradation risks.
  • Collaborate with Dev, QA, and infrastructureteams to identify and reduce toil and failure points.
  • Support Kubernetes workloads and assist introubleshooting cluster-level issues.
  • Work across AWS and Azure environments forincident containment and recovery.
  • Maintain and improve runbooks, playbooks, andincident response documentation.
  • Strong understanding of networking, security, anddistributed systems.
  • Excellent communication skills for cross-teamcollaboration and post-mortem documentation.
  • Experience with Helm, ArgoCD, or GitOpsworkflows.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

iLink Digital • Chennai

On-site
INR 1,200,000 - 1,800,000
DevOps Engineer
DevOps Engineer

iLink Digital • Chennai

On-site
INR 1,000,000 - 1,500,000
DevOps Engineer
DevOps Engineer

iLink Digital • Chennai District

On-site
INR 1,200,000 - 2,400,000
Devops/SRE
Devops/SRE

iLink Digital • Chennai District

On-site
INR 900,000 - 1,400,000
Devops/SRE
Devops/SRE

iLink Digital • Chennai District

On-site
INR 1,200,000 - 1,800,000
Senior Software Engineer - Backend
Senior Software Engineer - Backend

iLink Digital • Pune District

On-site
INR 2,500,000 - 5,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000