Site Reliability Engineer

SourcingXPress

Maharashtra

On-site

INR 700,000 - 1,800,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EOV Digital Solutions Pvt Ltd in India seeks an experienced Site Reliability Engineer (SRE) to design and scale AI-powered reliability across cloud-native AWS environments. You will own end-to-end reliability strategy, automation, and proactive engineering to ensure high availability and performance.

You will implement Terraform/CloudFormation, build dashboards with OpenTelemetry, Grafana, and Datadog, and apply SRE best practices including SLIs, SLOs and RCAs, collaborating with DevOps,

Qualifications

  • 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
  • Hands-on experience with AWS services (EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3).
  • Experience with OpenTelemetry, Grafana, Datadog, Prometheus or similar monitoring platforms.
  • Strong Linux administration, networking, and performance troubleshooting skills.
  • Experience with Infrastructure as Code: Terraform or CloudFormation.
  • Scripting in Python, Bash, or Go.
  • Experience with Kubernetes and containerized workloads.
  • Knowledge of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI).
  • Experience leading incident management, production support, and RCA processes.
  • Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.

Responsibilities

  • Design, implement, and manage highly available cloud infrastructure on AWS.
  • Build and maintain a complete observability platform using Open Telemetry, Grafana, Datadog, CloudWatch.
  • Implement AIOps capabilities including LLM-assisted incident triage and ML-driven forecasting.
  • Lead production incident management, on-call response, postmortems, and RCA.
  • Automate workflows using Terraform/CloudFormation and CI/CD pipelines.
  • Drive capacity planning, cost optimization, and rightsizing of infrastructure.
  • Build dashboards, SLOs, SLIs, and error budgets for reliability.
  • Develop automation scripts in Python, Bash, or Go to reduce manual tasks.
  • Monitor health of applications and infrastructure; ensure high uptime and performance.
  • Collaborate across Development, DevOps, Security, and Platform Engineering teams across time zones.
  • Establish best practices for monitoring, incident response, disaster recovery, and resilience.

Skills

SRE Experience
AWS Expertise
Linux Admin
CI/CD

Tools

Open Telemetry
Grafana
Datadog
Prometheus
Terraform
CloudFormation
Kubernetes
GitHub Actions
Jenkins
GitLab CI

Job description

Site Reliability Engineer (SRE) – AI & Cloud Infrastructure

We are looking for an experienced Site Reliability Engineer (SRE) to build and scale AI-powered reliability capabilities from the ground up. In this role, you will drive modern observability, automation, and cloud reliability initiatives while leveraging AI/ML for incident management, forecasting, and infrastructure optimization.

You will own the end-to-end reliability strategy across cloud-native AWS environments, enabling high availability, performance, and operational excellence through automation, intelligent monitoring, and proactive engineering.

Company: EOV Digital Solutions Pvt Ltd

Website: Visit Website

LinkedIn: Visit LinkedIn

Business Type: Small/Medium Business

Company Type: Product & Service

Business Model: B2B

Funding Stage: Private Equity

Industry: Information Technology

Salary Range: ₹ 7-18 Lacs PA

Key Responsibilities
  • Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS.
  • Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools.
  • Implement AIOps capabilities, including:
    • LLM-assisted incident triage
    • AI-powered root cause analysis
    • ML-driven forecasting and anomaly detection
    • Intelligent alert correlation and noise reduction
  • Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA).
  • Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines.
  • Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization.
  • Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.
  • Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks.
  • Monitor application and infrastructure health while ensuring high uptime and service performance.
  • Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones.
  • Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering.
  • Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues.
Required Skills & Qualifications
  • 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
  • Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling.
  • Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms.
  • Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting.
  • Experience with Infrastructure as Code using Terraform or CloudFormation.
  • Proficiency in scripting using Python, Bash, or Go.
  • Experience with Kubernetes and containerized workloads.
  • Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.).
  • Experience leading incident management, production support, and RCA processes.
  • Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.
  • Experience implementing monitoring, logging, alerting, and observability frameworks.
  • Strong analytical, troubleshooting, and communication skills.
Preferred Qualifications
  • Experience building or implementing AIOps solutions.
  • Exposure to Large Language Models (LLMs) for operational automation.
  • Experience with machine learning-based forecasting or anomaly detection.
  • Hands-on experience administering Adobe Experience Manager (AEM).
  • Experience managing Cloudflare CDN, WAF, DNS, and caching strategies.
  • Knowledge of FinOps, cloud cost optimization, and capacity planning.
  • AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Ahmedabad District

Hybrid
INR 400,000 - 700,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Logikality • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Opportunity to build a platform from an early stage
Direct exposure to leadership and product strategy
Significant growth opportunities for high performers
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

On-site
INR 900,000 - 1,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hilabs • Bengaluru

On-site
INR 2,200,000 - 3,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
AWS Site Reliability Engineer (SRE)
AWS Site Reliability Engineer (SRE)

Zensar • Hyderabad, Pune District

Hybrid
INR 1,500,000 - 2,100,000