Site Reliability Engineer

SourcingXPress

Maharashtra

On-site

INR 700,000 - 1,800,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EOV Digital Solutions Pvt Ltd in India seeks an experienced Site Reliability Engineer (SRE) to design and scale AI-powered reliability across cloud-native AWS environments. You will own end-to-end reliability strategy, automation, and proactive engineering to ensure high availability and performance.

You will implement Terraform/CloudFormation, build dashboards with OpenTelemetry, Grafana, and Datadog, and apply SRE best practices including SLIs, SLOs and RCAs, collaborating with DevOps,

Qualifications

  • 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
  • Hands-on experience with AWS services (EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3).
  • Experience with OpenTelemetry, Grafana, Datadog, Prometheus or similar monitoring platforms.
  • Strong Linux administration, networking, and performance troubleshooting skills.
  • Experience with Infrastructure as Code: Terraform or CloudFormation.
  • Scripting in Python, Bash, or Go.
  • Experience with Kubernetes and containerized workloads.
  • Knowledge of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI).
  • Experience leading incident management, production support, and RCA processes.
  • Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.

Responsibilities

  • Design, implement, and manage highly available cloud infrastructure on AWS.
  • Build and maintain a complete observability platform using Open Telemetry, Grafana, Datadog, CloudWatch.
  • Implement AIOps capabilities including LLM-assisted incident triage and ML-driven forecasting.
  • Lead production incident management, on-call response, postmortems, and RCA.
  • Automate workflows using Terraform/CloudFormation and CI/CD pipelines.
  • Drive capacity planning, cost optimization, and rightsizing of infrastructure.
  • Build dashboards, SLOs, SLIs, and error budgets for reliability.
  • Develop automation scripts in Python, Bash, or Go to reduce manual tasks.
  • Monitor health of applications and infrastructure; ensure high uptime and performance.
  • Collaborate across Development, DevOps, Security, and Platform Engineering teams across time zones.
  • Establish best practices for monitoring, incident response, disaster recovery, and resilience.

Skills

SRE Experience
AWS Expertise
Linux Admin
CI/CD

Tools

Open Telemetry
Grafana
Datadog
Prometheus
Terraform
CloudFormation
Kubernetes
GitHub Actions
Jenkins
GitLab CI

Job description

Site Reliability Engineer (SRE) – AI & Cloud Infrastructure

We are looking for an experienced Site Reliability Engineer (SRE) to build and scale AI-powered reliability capabilities from the ground up. In this role, you will drive modern observability, automation, and cloud reliability initiatives while leveraging AI/ML for incident management, forecasting, and infrastructure optimization.

You will own the end-to-end reliability strategy across cloud-native AWS environments, enabling high availability, performance, and operational excellence through automation, intelligent monitoring, and proactive engineering.

Company: EOV Digital Solutions Pvt Ltd

Website: Visit Website

LinkedIn: Visit LinkedIn

Business Type: Small/Medium Business

Company Type: Product & Service

Business Model: B2B

Funding Stage: Private Equity

Industry: Information Technology

Salary Range: ₹ 7-18 Lacs PA

Key Responsibilities
  • Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS.
  • Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools.
  • Implement AIOps capabilities, including:
    • LLM-assisted incident triage
    • AI-powered root cause analysis
    • ML-driven forecasting and anomaly detection
    • Intelligent alert correlation and noise reduction
  • Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA).
  • Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines.
  • Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization.
  • Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.
  • Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks.
  • Monitor application and infrastructure health while ensuring high uptime and service performance.
  • Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones.
  • Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering.
  • Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues.
Required Skills & Qualifications
  • 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
  • Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling.
  • Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms.
  • Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting.
  • Experience with Infrastructure as Code using Terraform or CloudFormation.
  • Proficiency in scripting using Python, Bash, or Go.
  • Experience with Kubernetes and containerized workloads.
  • Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.).
  • Experience leading incident management, production support, and RCA processes.
  • Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.
  • Experience implementing monitoring, logging, alerting, and observability frameworks.
  • Strong analytical, troubleshooting, and communication skills.
Preferred Qualifications
  • Experience building or implementing AIOps solutions.
  • Exposure to Large Language Models (LLMs) for operational automation.
  • Experience with machine learning-based forecasting or anomaly detection.
  • Hands-on experience administering Adobe Experience Manager (AEM).
  • Experience managing Cloudflare CDN, WAF, DNS, and caching strategies.
  • Knowledge of FinOps, cloud cost optimization, and capacity planning.
  • AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer (SRE) – AI & Cloud Infrastructure
Site Reliability Engineer (SRE) – AI & Cloud Infrastructure

TeamPlus Staffing Solution Pvt Ltd • Pune District

On-site
INR 1,800,000 - 3,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Logikality • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Opportunity to build a platform from an early stage
Direct exposure to leadership and product strategy
Significant growth opportunities for high performers
AWS Site Reliability Engineer (SRE)
AWS Site Reliability Engineer (SRE)

Zensar • Hyderabad, Pune District

Hybrid
INR 1,500,000 - 2,100,000
Senior Site Reliability Engineer - Cloud Infrastructure
Senior Site Reliability Engineer - Cloud Infrastructure

WITS Innovation Lab • Chandigarh

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

Hybrid
INR 900,000 - 1,400,000
Senior Site Reliability Engineer (SRE) – AWS
Senior Site Reliability Engineer (SRE) – AWS

Sailssoftware • Visakhapatnam

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000