VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd.

Hinoba-an

On-site

PHP 893,000 - 1,674,000

Full time

12 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

E4 Software Services Pvt Ltd. is seeking an experienced SRE/DevOps professional to monitor, manage and improve production environments. You will focus on reliability, observability and automation to reduce toil.

Responsibilities include incident handling, RCA, root cause actions, and collaboration with development and infrastructure teams to boost resilience. Knowledge of cloud platforms (AWS/Azure/GCP), Kubernetes, Docker, and CI/CD is essential.

Qualifications

  • 4+ years in SRE, Production Support, DevOps or Cloud Operations.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with monitoring and observability platforms.
  • Hands-on experience with Splunk, ELK, Dynatrace, Datadog, Prometheus or Grafana.
  • Strong scripting using Python, Bash or Shell.
  • Experience with incident management, problem management and RCA.
  • Good understanding of cloud platforms such as AWS, Azure or GCP.
  • Knowledge of networking concepts including TCP/IP, DNS and load balancing.
  • Experience with CI/CD and DevOps practices.
  • Strong production troubleshooting and operational support experience.

Responsibilities

  • Monitor and maintain the availability, performance and reliability of production environments.
  • Design and manage monitoring, logging, alerting and observability solutions.
  • Analyze application and infrastructure logs to identify performance and reliability issues.
  • Handle production incidents, service requests and operational escalations.
  • Perform root cause analysis and implement permanent corrective actions.
  • Develop automation to reduce repetitive operational activities and manual intervention.
  • Define and monitor reliability indicators, SLIs, SLOs and operational KPIs.
  • Support cloud environments across AWS, Azure and/or GCP.
  • Troubleshoot Linux/Unix, networking, application and infrastructure issues.
  • Support Kubernetes, Docker and CI/CD environments where required.
  • Maintain operational runbooks and improve incident response processes.
  • Work with development and infrastructure teams to improve system resilience.
  • Participate in on-call and production support activities.

Skills

SRE Experience
Linux Administration
Monitoring/Observability
Monitoring Tools
Scripting
Incident RCA
Cloud Platforms
Networking
CI/CD
Operational Support

Tools

Kubernetes
Docker
Apache/Tomcat
ITSM tools (ServiceNow/Jira)
Ansible/Chef/Puppet
SAP support
OpenTelemetry/Moogsoft/Rundeck
Prometheus/Grafana

Job description

  • Monitor and maintain the availability, performance and reliability of production environments.
  • Design and manage monitoring, logging, alerting and observability solutions.
  • Analyze application and infrastructure logs to identify performance and reliability issues.
  • Handle production incidents, service requests and operational escalations.
  • Perform root cause analysis and implement permanent corrective actions.
  • Develop automation to reduce repetitive operational activities and manual intervention.
  • Define and monitor reliability indicators, SLIs, SLOs and operational KPIs.
  • Support cloud environments across AWS, Azure and/or GCP.
  • Troubleshoot Linux/Unix, networking, application and infrastructure issues.
  • Support Kubernetes, Docker and CI/CD environments where required.
  • Maintain operational runbooks and improve incident response processes.
  • Work with development and infrastructure teams to improve system resilience.
  • Participate in on-call and production support activities.
Mandatory Skills
  • 4+ years of experience in SRE, Production Support, DevOps or Cloud Operations.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience with monitoring and observability platforms.
  • Hands-on experience with Splunk, ELK, Dynatrace, Datadog, Prometheus or Grafana.
  • Strong scripting skills using Python, Bash or Shell.
  • Experience with incident management, problem management and RCA.
  • Good understanding of cloud platforms such as AWS, Azure or GCP.
  • Knowledge of networking concepts including TCP/IP, DNS and load balancing.
  • Experience with CI/CD and DevOps practices.
  • Strong production troubleshooting and operational support experience.
Good-to-have Skills
  • OpenTelemetry, Moogsoft or Rundeck.
  • Kubernetes and Docker.
  • Apache, Tomcat or enterprise application platforms.
  • ServiceNow, Jira or other ITSM tools.
  • Ansible, Chef or Puppet.
  • SAP application/platform support.
  • Experience with SLI/SLO and reliability engineering practices.
  • Experience with automated remediation and AIOps.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Lead Application Support Engineer
Lead Application Support Engineer

Morgan McKinley • Taguig

On-site
PHP 1,200,000 - 1,600,000
VS01699 - Senior DevOps & Cloud Engineer
VS01699 - Senior DevOps & Cloud Engineer

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 800,000 - 1,400,000
Reliability Engineer: SRE, Observability & Cloud Ops
Reliability Engineer: SRE, Observability & Cloud Ops

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA PHILIPPINES INC. • Philippines

On-site
PHP 900,000 - 1,350,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8inc • Manila

On-site
PHP 1,200,000 - 1,600,000
Site Reliability / Cloud Platform Engineer
Site Reliability / Cloud Platform Engineer

Global Recruitment and Consultancy OPC • Cebu City

On-site
PHP 1,200,000 - 2,400,000
Senior Production Engineer
Senior Production Engineer

863 Parameta Solutions (Singapore) Pte. Limited • Taguig

On-site
PHP 558,000 - 781,200
Site Reliability Engineer
Site Reliability Engineer

Alsons/AWS Information Systems Inc. • Cebu City

Hybrid
PHP 600,000 - 1,000,000
Associate Site Reliability Engineer
Associate Site Reliability Engineer

Railway Corp • Mexico

On-site
PHP 5,474,000 - 7,908,000