SRE - AWS DevOPS Engineer

Prowess Publishing

Hyderabad

On-site

INR 900,000 - 1,500,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Prowess Publishing in Hyderabad is seeking a hands-on Site Reliability Engineer (SRE) / AWS Cloud Operations Engineer with 3–4 years of experience to own production infrastructure, cloud operations, incident management and reliability.

The role emphasizes troubleshooting production issues, automation, and building observability, with a focus on AWS, Linux, Kubernetes/Docker, and CI/CD. This is an in-office position.

Qualifications

  • 3–4 years of relevant experience in SRE, AWS DevOps, Cloud Operations or Production Engineering.
  • Strong hands-on AWS Cloud experience.
  • Strong Production Support / Production Operations experience.
  • Excellent Linux troubleshooting skills.
  • Hands-on experience with Kubernetes / EKS and Docker.
  • Good understanding of AWS networking – VPC, subnets, security groups, load balancers, routing, DNS, etc.
  • Experience with monitoring and alerting tools.
  • Strong troubleshooting, incident management and RCA skills.
  • Experience with Bash/Shell or Python scripting.
  • Good understanding of CI/CD concepts and tools.
  • Ability to independently investigate and resolve production issues.

Responsibilities

  • Own and support production AWS infrastructure and cloud environments.
  • Monitor applications, infrastructure and services; identify issues proactively.
  • Troubleshoot and resolve production incidents within defined SLAs.
  • Perform root cause analysis (RCA) for recurring production issues and implement permanent fixes.
  • Manage AWS services: EC2, ECS/EKS, S3, IAM, VPC, ELB/ALB, CloudWatch.
  • Work with Linux environments: system processes, logs, networking and performance.
  • Support and troubleshoot Kubernetes/Docker workloads.
  • Implement infrastructure automation using Terraform / IaC.
  • Build and maintain monitoring and observability using CloudWatch, Prometheus, Grafana, Datadog, Splunk.
  • Automate repetitive operational activities using Python, Bash or Shell scripting.
  • Support CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD.
  • Participate in incident response, problem management and post-incident reviews.
  • Collaborate with Development, QA and other teams to improve reliability and deployment processes.
  • Identify opportunities for automation and capacity/performance improvements.
  • Participate in on-call / production support activities as required.

Skills

SRE
AWS Cloud
Linux troubleshooting
Kubernetes
Docker
RCA
CI/CD
Python scripting
Bash scripting
Incident management
On-call

Tools

Terraform / IaC
CloudWatch
Prometheus
Grafana
Datadog
Splunk
Jenkins
GitHub Actions
GitLab CI/CD
ELK / OpenSearch

Job description

Site Reliability Engineer (SRE) AWS DevOPS

Experience: 34 Years
Location: Hyderabad
Work Mode: Work from Office

Job Summary

We are looking for a hands-on Site Reliability Engineer (SRE) / AWS Cloud Operations Engineer with 3–4 years of experience to take ownership of production infrastructure, cloud operations, incident management and reliability.

The ideal candidate should be comfortable working in a fast-paced environment, troubleshooting production issues independently, monitoring critical systems and implementing automation to improve system availability and operational efficiency.

This is not a generic DevOps role. Strong hands-on experience in AWS, Linux, production support, troubleshooting and incident management is essential.

Key Responsibilities
  • Own and support production AWS infrastructure and cloud environments.
  • Monitor applications, infrastructure and services and proactively identify performance or availability issues.
  • Troubleshoot and resolve production incidents within defined SLAs.
  • Perform root cause analysis (RCA) for recurring production issues and implement permanent fixes.
  • Manage AWS services including EC2, ECS/EKS, S3, IAM, VPC, ELB/ALB, CloudWatch and related services.
  • Work with Linux environments, system processes, logs, networking and performance troubleshooting.
  • Support and troubleshoot Kubernetes/Docker based workloads.
  • Implement infrastructure automation using Terraform / Infrastructure as Code.
  • Build and maintain monitoring, alerting and observability using CloudWatch, Prometheus, Grafana, Datadog, Splunk or equivalent tools.
  • Automate repetitive operational activities using Python, Bash or Shell scripting.
  • Support and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD or similar tools.
  • Participate in incident response, problem management and post-incident reviews.
  • Collaborate with Development, QA and other engineering teams to improve application reliability and deployment processes.
  • Identify opportunities for automation, capacity improvement, performance optimization and reduction of operational toil.
  • Participate in on-call / production support activities as required.
Mandatory Skills
  • 3–4 years of relevant experience in SRE, AWS DevOps, Cloud Operations or Production Engineering.
  • Strong hands-on AWS Cloud experience.
  • Strong Production Support / Production Operations experience.
  • Excellent Linux troubleshooting skills.
  • Hands-on experience with Kubernetes / EKS and Docker.
  • Good understanding of AWS networking – VPC, subnets, security groups, load balancers, routing, DNS, etc.
  • Experience with monitoring and alerting tools.
  • Strong troubleshooting, incident management and RCA skills.
  • Experience with Bash/Shell or Python scripting.
  • Good understanding of CI/CD concepts and tools.
  • Ability to independently investigate and resolve production issues.
Good to Have
  • Terraform / CloudFormation
  • Prometheus / Grafana / Datadog
  • AWS ECS / Fargate
  • Jenkins / GitHub Actions / GitLab
  • ELK / OpenSearch / Splunk
  • Infrastructure as Code
  • Auto Scaling and high-availability architecture
  • Performance and capacity monitoring
  • ITIL / Incident & Problem Management
  • Experience supporting microservices environments
What We Are Looking For

We are particularly interested in candidates who demonstrate:

Production Ownership | Strong AWS | Linux Troubleshooting | Incident Management | RCA | Kubernetes | Monitoring | Automation


Candidates who have primarily worked on CI/CD pipeline development without significant production-support ownership may not be suitable for this position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Lyzr AI • Bengaluru

Hybrid
INR 1,000,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

Hybrid
INR 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Unified Consultancy Services • Karnataka

Hybrid
INR 1,500,000 - 3,000,000
Senior Site Reliability Engineer - Cloud Infrastructure
Senior Site Reliability Engineer - Cloud Infrastructure

WITS Innovation Lab • Chandigarh

On-site
INR 1,800,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
SRE+AWS Devops
SRE+AWS Devops

Virtusa • Bengaluru Urban

On-site
INR 1,500,000 - 2,500,000