VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd.

India

On-site

INR 2,000,000 - 4,000,000

Full time

19 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

E4 Software Services Pvt Ltd. in India is seeking an experienced SRE to monitor and maintain production environments, design observability solutions and manage incidents.

You will work with Linux, cloud platforms, and containers, implement automation, and drive reliability metrics and RCA for continuous improvement.

Qualifications

  • 4+ years of experience in SRE, Production Support, DevOps or Cloud Operations.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with monitoring and observability platforms.
  • Hands-on experience with Splunk, ELK, Dynatrace, Datadog, Prometheus or Grafana.
  • Strong scripting skills using Python, Bash or Shell.
  • Experience with incident management, problem management and RCA.
  • Good understanding of cloud platforms such as AWS, Azure or GCP.
  • Knowledge of networking concepts including TCP/IP, DNS and load balancing.
  • Experience with CI/CD and DevOps practices.
  • Strong production troubleshooting and operational support experience.

Responsibilities

  • Monitor and maintain the availability, performance and reliability of production environments.
  • Design and manage monitoring, logging, alerting and observability solutions.
  • Analyze application and infrastructure logs to identify performance and reliability issues.
  • Handle production incidents, service requests and operational escalations.
  • Perform root cause analysis and implement permanent corrective actions.
  • Develop automation to reduce repetitive operational activities and manual intervention.
  • Define and monitor reliability indicators, SLIs, SLOs and operational KPIs.
  • Support cloud environments across AWS, Azure and/or GCP.
  • Troubleshoot Linux/Unix, networking, application and infrastructure issues.
  • Support Kubernetes, Docker and CI/CD environments where required.
  • Maintain operational runbooks and improve incident response processes.
  • Work with development and infrastructure teams to improve system resilience.
  • Participate in on-call and production support activities.

Skills

Linux/Unix admin
Observability
SRE/DevOps mindset
Incident management
Cloud platforms AWS/Azure/GCP
Networking concepts
Scripting (Python/Bash)
CI/CD practices

Tools

Splunk
ELK
Dynatrace
Datadog
Prometheus
Grafana
Kubernetes
Docker
CI/CD tooling

Job description

  • Monitor and maintain the availability, performance and reliability of production environments.
  • Design and manage monitoring, logging, alerting and observability solutions.
  • Analyze application and infrastructure logs to identify performance and reliability issues.
  • Handle production incidents, service requests and operational escalations.
  • Perform root cause analysis and implement permanent corrective actions.
  • Develop automation to reduce repetitive operational activities and manual intervention.
  • Define and monitor reliability indicators, SLIs, SLOs and operational KPIs.
  • Support cloud environments across AWS, Azure and/or GCP.
  • Troubleshoot Linux/Unix, networking, application and infrastructure issues.
  • Support Kubernetes, Docker and CI/CD environments where required.
  • Maintain operational runbooks and improve incident response processes.
  • Work with development and infrastructure teams to improve system resilience.
  • Participate in on-call and production support activities.
Mandatory Skills
  • 4+ years of experience in SRE, Production Support, DevOps or Cloud Operations.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with monitoring and observability platforms.
  • Hands-on experience with Splunk, ELK, Dynatrace, Datadog, Prometheus or Grafana.
  • Strong scripting skills using Python, Bash or Shell.
  • Experience with incident management, problem management and RCA.
  • Good understanding of cloud platforms such as AWS, Azure or GCP.
  • Knowledge of networking concepts including TCP/IP, DNS and load balancing.
  • Experience with CI/CD and DevOps practices.
  • Strong production troubleshooting and operational support experience.
Good-to-have Skills
  • OpenTelemetry, Moogsoft or Rundeck.
  • Kubernetes and Docker.
  • Apache, Tomcat or enterprise application platforms.
  • ServiceNow, Jira or other ITSM tools.
  • Ansible, Chef or Puppet.
  • SAP application/platform support.
  • Experience with SLI/SLO and reliability engineering practices.
  • Experience with automated remediation and AIOps.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Expert
SRE Expert

HCLTech • Bengaluru

On-site
INR 1,500,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
SRE- Production Support
SRE- Production Support

Cloudxtreme • Hyderabad

On-site
INR 2,400,000 - 3,600,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
AWS SRE Professional
AWS SRE Professional

Infosys • Bengaluru

On-site
INR 900,000 - 1,300,000
SRE Lead
SRE Lead

Bounteous • Gurugram District

On-site
INR 1,200,000 - 2,400,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Lead Support Analyst - Shared Services and Production Management , Information Technology
Lead Support Analyst - Shared Services and Production Management , Information Technology

CLSA • Pune District

On-site
INR 1,500,000 - 2,800,000