Site Reliability Engineer

PeoplePlusTech Inc.

Metro Manila

Hybrid

PHP 700,000 - 1,100,000

Full time

14 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

PeoplePlusTech Inc. is seeking a Site Reliability Engineer (SRE) to join our growing technology team.

The ideal candidate will maintain the reliability, availability, scalability, and performance of critical applications and infrastructure, blending software engineering with systems administration to build automated, highly available platforms. The successful candidate will collaborate with development, infrastructure, security, and operations teams to drive automation, improve reliability, and

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • Minimum of 3 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Systems Engineering, or Infrastructure Operations.
  • Experience supporting production environments with high availability and uptime requirements.
  • Strong understanding of Linux and/or Windows server administration.
  • Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Experience with containerization technologies such as Docker and orchestration platforms such as Kubernetes.
  • Hands-on experience with CI/CD tools such as Azure DevOps, GitHub Actions, Jenkins, GitLab CI/CD, or similar.
  • Experience with infrastructure-as-code tools such as Terraform, CloudFormation, or Ansible.
  • Familiarity with monitoring and observability platforms such as Prometheus, Grafana, Datadog, New Relic, Dynatrace, Splunk, or ELK Stack.
  • Strong scripting and automation skills using Python, Bash, PowerShell, or similar languages.
  • Experience in troubleshooting complex production incidents and conducting root cause analysis.
  • Knowledge of networking concepts, DNS, load balancing, firewalls, and security best practices.

Responsibilities

  • Design, implement, and maintain highly available and scalable infrastructure and application environments.
  • Monitor system health, application performance, and service availability using industry-standard monitoring and observability tools.
  • Develop and maintain automation scripts, tools, and workflows to improve operational efficiency.
  • Manage incident response, troubleshooting, root cause analysis (RCA), and post-incident reviews.
  • Collaborate with development teams to improve application reliability, deployment processes, and operational readiness.
  • Establish and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
  • Implement infrastructure-as-code (IaC) solutions for consistent and repeatable deployments.
  • Support CI/CD pipelines and deployment automation initiatives.
  • Perform capacity planning, performance tuning, and proactive system optimization.
  • Ensure compliance with security, governance, and operational best practices.
  • Create and maintain technical documentation, operational runbooks, and disaster recovery procedures.
  • Participate in on-call support and incident management activities as required.

Skills

Linux
Windows Server
AWS
Azure
GCP
Docker
Kubernetes
CI/CD
Terraform
Ansible
Python
Bash
PowerShell
Networking fundamentals
Prometheus
Grafana
Security best practices
Incident management
Root cause analysis

Education

Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field

Job description

Job Title: Site Reliability Engineer (SRE)

Work Setup: Hybrid (3 Days Onsite, 2 Days Remote)

Employment Type: Full-Time

Job Summary

We are seeking a highly skilled Site Reliability Engineer (SRE) to join our growing technology team. The ideal candidate will be responsible for maintaining the reliability, availability, scalability, and performance of critical applications and infrastructure. This role combines software engineering and systems administration principles to build and operate resilient, automated, and highly available platforms.

The successful candidate will work closely with development, infrastructure, security, and operations teams to improve system reliability, optimize performance, and enhance operational efficiency through automation and continuous improvement initiatives.

Key Responsibilities

  • Design, implement, and maintain highly available and scalable infrastructure and application environments.
  • Monitor system health, application performance, and service availability using industry-standard monitoring and observability tools.
  • Develop and maintain automation scripts, tools, and workflows to improve operational efficiency.
  • Manage incident response, troubleshooting, root cause analysis (RCA), and post-incident reviews.
  • Collaborate with development teams to improve application reliability, deployment processes, and operational readiness.
  • Establish and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
  • Implement infrastructure-as-code (IaC) solutions for consistent and repeatable deployments.
  • Support CI/CD pipelines and deployment automation initiatives.
  • Perform capacity planning, performance tuning, and proactive system optimization.
  • Ensure compliance with security, governance, and operational best practices.
  • Create and maintain technical documentation, operational runbooks, and disaster recovery procedures.
  • Participate in on-call support and incident management activities as required.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • Minimum of 3 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Systems Engineering, or Infrastructure Operations.
  • Experience supporting production environments with high availability and uptime requirements.
  • Strong understanding of Linux and/or Windows server administration.
  • Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Experience with containerization technologies such as Docker and orchestration platforms such as Kubernetes.
  • Hands-on experience with CI/CD tools such as Azure DevOps, GitHub Actions, Jenkins, GitLab CI/CD, or similar.
  • Experience with infrastructure-as-code tools such as Terraform, CloudFormation, or Ansible.
  • Familiarity with monitoring and observability platforms such as Prometheus, Grafana, Datadog, New Relic, Dynatrace, Splunk, or ELK Stack.
  • Strong scripting and automation skills using Python, Bash, PowerShell, or similar languages.
  • Experience in troubleshooting complex production incidents and conducting root cause analysis.
  • Knowledge of networking concepts, DNS, load balancing, firewalls, and security best practices.

Preferred Qualifications

  • Experience working in enterprise or cloud-native environments.
  • Familiarity with microservices architecture and distributed systems.
  • Experience implementing reliability engineering practices, including SLOs, SLIs, and error budgets.
  • Cloud certifications (AWS, Azure, or GCP) are an advantage.
  • Experience with security and compliance frameworks.
  • Knowledge of disaster recovery and business continuity planning.

Technical Skills

  • Terraform, Ansible, CloudFormation
  • Monitoring & Observability Tools
  • Python, Bash, PowerShell
  • Networking & Security Fundamentals
  • Incident Management & Root Cause Analysis
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

IDEMIA • Philippines

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA PHILIPPINES INC. • Philippines

On-site
PHP 900,000 - 1,350,000
DevOps Engineer / Site Reliability Engineer (SRE)
DevOps Engineer / Site Reliability Engineer (SRE)

Concentrix • Mexico

On-site
PHP 1,000,000 - 1,600,000
Site Reliability Engineer
Site Reliability Engineer

V2 Solutions • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Pyramid Consulting, Inc • Mexico

On-site
PHP 5,618,000 - 8,115,000
Site Reliability Engineer
Site Reliability Engineer

AgileEngine • Mexico

Hybrid
PHP 8,734,000 - 13,100,000
Professional growth
Competitive USD-based pay
Exciting projects
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

EPAM Systems • Mexico

Hybrid
PHP 900,000 - 1,300,000
Healthcare benefits
Paid time off and sick leave
Learning & development programs
+2
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000