VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd.

Bengaluru

On-site

INR 1,800,000 - 2,400,000

Full time

30 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

E4 Software Services Pvt Ltd. in Bengaluru seeks an experienced Site Reliability Engineer to maintain production environments, ensuring high availability and reliability.

You will design observability stacks and automate incident response, collaborating with development and infrastructure teams. You will monitor logs, configure SLIs/SLOs, support AWS/Azure/GCP, and troubleshoot Linux, networking, Kubernetes and CI/CD pipelines.

Qualifications

  • 4+ years in SRE, production support, DevOps or cloud operations.
  • Strong Linux/Unix administration and troubleshooting.
  • Experience with monitoring and observability platforms.
  • Hands-on experience with Splunk, ELK, Dynatrace, Datadog, Prometheus or Grafana.
  • Strong scripting with Python, Bash or Shell.
  • Experience with incident management and RCA.
  • Knowledge of cloud platforms (AWS, Azure or GCP).
  • Networking concepts including TCP/IP, DNS and load balancing.
  • Experience with CI/CD and DevOps practices.
  • Production troubleshooting and operational support.

Responsibilities

  • Monitor and maintain the availability, performance and reliability of production environments.
  • Design and manage monitoring, logging, alerting and observability solutions.
  • Analyze application and infrastructure logs to identify performance and reliability issues.
  • Handle production incidents, service requests and operational escalations.
  • Perform root cause analysis and implement permanent corrective actions.
  • Develop automation to reduce repetitive operational activities and manual intervention.
  • Define and monitor reliability indicators, SLIs, SLOs and operational KPIs.
  • Support cloud environments across AWS, Azure and/or GCP.
  • Troubleshoot Linux/Unix, networking, application and infrastructure issues.
  • Support Kubernetes, Docker and CI/CD environments where required.
  • Maintain operational runbooks and improve incident response processes.
  • Work with development and infrastructure teams to improve system resilience.
  • Participate in on-call and production support activities.

Skills

SRE/DevOps
Linux/Unix admin
Monitoring/Observability
Python scripting
Incident & RCA
Cloud platforms
Networking basics
CI/CD
Troubleshooting

Tools

Splunk
ELK Stack
Dynatrace
Datadog
Prometheus
Grafana
Kubernetes
Docker
Ansible
Jira

Job description

  • Monitor and maintain the availability, performance and reliability of production environments.
  • Design and manage monitoring, logging, alerting and observability solutions.
  • Analyze application and infrastructure logs to identify performance and reliability issues.
  • Handle production incidents, service requests and operational escalations.
  • Perform root cause analysis and implement permanent corrective actions.
  • Develop automation to reduce repetitive operational activities and manual intervention.
  • Define and monitor reliability indicators, SLIs, SLOs and operational KPIs.
  • Support cloud environments across AWS, Azure and/or GCP.
  • Troubleshoot Linux/Unix, networking, application and infrastructure issues.
  • Support Kubernetes, Docker and CI/CD environments where required.
  • Maintain operational runbooks and improve incident response processes.
  • Work with development and infrastructure teams to improve system resilience.
  • Participate in on-call and production support activities.
Mandatory Skills
  • 4+ years of experience in SRE, Production Support, DevOps or Cloud Operations.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with monitoring and observability platforms.
  • Hands-on experience with Splunk, ELK, Dynatrace, Datadog, Prometheus or Grafana.
  • Strong scripting skills using Python, Bash or Shell.
  • Experience with incident management, problem management and RCA.
  • Good understanding of cloud platforms such as AWS, Azure or GCP.
  • Knowledge of networking concepts including TCP/IP, DNS and load balancing.
  • Experience with CI/CD and DevOps practices.
  • Strong production troubleshooting and operational support experience.
Good-to-have Skills
  • OpenTelemetry, Moogsoft or Rundeck.
  • Kubernetes and Docker.
  • Apache, Tomcat or enterprise application platforms.
  • ServiceNow, Jira or other ITSM tools.
  • Ansible, Chef or Puppet.
  • SAP application/platform support.
  • Experience with SLI/SLO and reliability engineering practices.
  • Experience with automated remediation and AIOps.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
SRE Developer
SRE Developer

Cloudxtreme • Hyderabad

On-site
INR 1,500,000 - 2,400,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Application SRE
Application SRE

Cloudxtreme • Pune District

On-site
INR 1,200,000 - 1,800,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
AWS SRE Professional
AWS SRE Professional

Infosys • Bengaluru

On-site
INR 900,000 - 1,300,000
SRE Engineer
SRE Engineer

ConsultBae India Private limited • India

On-site
INR 1,200,000 - 2,000,000
SRE Lead
SRE Lead

3across • Bengaluru

Hybrid
INR 1,500,000 - 2,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000