Site Reliability Engineer

Procter & Gamble

Manila

On-site

PHP 900,000 - 1,500,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

P&G Manila is seeking a Site Reliability Engineer for the Warehousing IT Operations - Incident Response Team. You will lead incident response efforts and work with software, DevOps, and site customers to restore services quickly and reliably.

You will design, implement, and automate monitoring, incident management, and resilience improvements to ensure uptime and scalable performance across our systems. You will participate in post-incident reviews to drive continuous improvement.

Qualifications

  • Experience in incident response and post-incident reviews.
  • Strong focus on system reliability and monitoring.
  • Ability to collaborate with software, DevOps, and site customers.

Responsibilities

  • Lead incident response efforts, swiftly resolving critical incidents to minimize downtime and user impact.
  • Implement effective incident management processes with clear communication and documentation.
  • Conduct root cause analysis and drive preventive measures for continuous improvement.
  • Ensure high system availability through robust monitoring and automated responses.
  • Collaborate cross-functionally to design resilient systems and optimized architectures.
  • Configure and manage monitoring to provide real-time insights and actionable alerts.

Skills

Incident response
Monitoring
Reliability engineering
Root cause analysis
Automation
Cross-functional collaboration
DevOps practices

Tools

Monitoring tools
Incident management tools
Automation scripts

Job description

Job Location

MANILA NET PARK OFFICE

Job Description

At P&G, your career is built to scale with our iconic brands like Ariel ® , Safeguard ® , Tide ® , Head & Shoulders ® , Old Spice ® , and Vicks ®.

Overview of the job

As a Site Reliability Engineer (SRE) in the Warehousing IT Operations - Incident Response Team, you will be responsible for leading incident response efforts, ensuring swift and effective resolution of critical system issues. You will also play a critical role in ensuring the reliability, scalability, and performance of our systems and services. SRE combines software engineering and operations to build, maintain, and support highly available and efficient infrastructure. Your expertise in troubleshooting and root cause analysis will be essential in identifying and addressing the underlying causes of incidents. You will work closely with software engineers, DevOps teams, and other stakeholders to implement preventive measures and enhance system resilience. Collaborating with cross‑functional teams, you will design, implement, and automate robust systems, monitoring tools, and processes. With a strong focus on stability and uptime, you will proactively identify and resolve performance bottlenecks, optimize system architecture, and drive continuous improvement. Your keen eye for continuous improvement will also drive post-incident reviews and contribute to the creation of incident management best practices. By actively monitoring system health, responding to incidents in a timely manner, and implementing proactive measures, you will play a pivotal role in maintaining the stability and availability of our services, ensuring an exceptional user experience for our customers.

Your team

You will report directly to the Incident Response Engineering Leader within the Warehousing IT Operations team, who will provide guidance, support, and mentorship as you navigate your role. As a valued member of our dynamic Incident Response Team, you will collaborate closely with technically skilled professionals, including software engineers, DevOps specialists, Subject Matter Experts, and other SREs. In addition, you will have the opportunity to directly collaborate with our site customers and users, ensuring their needs and expectations are met through reliable and high-performing systems. Working within a cross-functional and collaborative environment, you will contribute to the success of our Incident Response team, which is dedicated to ensuring the reliability and availability of our site’s systems. Our Incident Response team fosters a culture of technical expertise, continuous learning, and knowledge sharing, where ideas are encouraged, and innovation is embraced.

How success looks like

Success as a Site Reliability Engineer (SRE) involves different areas of the role including incident response, monitoring and reliability, and effectively collaborating with customers and users, addressing their needs and expectations:

  • Incident Response: Swiftly respond to and resolve critical incidents, ensuring minimal impact on system availability and user experience while driving continuous improvement in incident management processes.
  • Reliability: Ensure high system availability and reliability through robust monitoring, optimization of system architecture, and cross-functional collaboration to design and implement resilient systems.
  • Monitoring: Implement comprehensive monitoring solutions to gain real-time insights into system performance, enabling proactive incident response and continuous improvement of system visibility and resource optimization.
  • Working with Customers/Users: Collaborate directly with customers and users to understand their needs, proactively address concerns, and provide exceptional customer support to ensure reliable and performant systems that meet their expectations.
Responsibilities of the role

Incident Response:

  • Lead incident response efforts, swiftly resolving critical incidents to minimize downtime and user impact.
  • Implement effective incident management processes, ensuring clear communication, coordination, and documentation.
  • Conduct root cause analysis, implementing preventive measures and driving continuous improvement.

Reliability:

  • Ensure high system availability through robust monitoring, alerting, and automated incident response systems.
  • Optimize system architecture and configurations for improved performance, scalability, and fault tolerance.
  • Collaborate cross-functionally to design and implement resilient systems using industry best practices.

Monitoring:

  • Implement comprehensive monitoring solutions, providing real-time insights into system performance and health.
  • Configure and manage monitoring tools, ensuring accurate and actionable alerts for proactive incident response.
  • Continuously evaluate and enhance monitoring strategies to improve system visibility and resource optimization.
Upskilling
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Procter & Gamble • Philippines

On-site
PHP 1,000,000 - 1,400,000
Performance bonus
Flexible work schedule / work from hom
Health insurance
+2
Site Reliability Engineer - Warehousing IT Operations
Site Reliability Engineer - Warehousing IT Operations

Procter & Gamble • Philippines

On-site
PHP 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Procter & Gamble • Philippines

On-site
PHP 1,200,000 - 1,500,000
Performance bonus (STAR)
Flexible work schedule with work-from-
Health insurance
+2
Site Reliability Engineer — Incidents, Monitoring & Scale
Site Reliability Engineer — Incidents, Monitoring & Scale

Procter & Gamble • Philippines

On-site
PHP 1,000,000 - 1,400,000
Performance bonus
Flexible work schedule / work from hom
Health insurance
+2
SRE Manager — Incident Response & Reliability Leader
SRE Manager — Incident Response & Reliability Leader

Procter & Gamble • Philippines

On-site
PHP 1,200,000 - 1,800,000
Site Reliability Engineer: Incident Response Lead
Site Reliability Engineer: Incident Response Lead

Procter & Gamble • Manila

On-site
PHP 900,000 - 1,500,000
Senior SRE Lead — Incident Response & Reliability (Remote)
Senior SRE Lead — Incident Response & Reliability (Remote)

Procter & Gamble • Philippines

On-site
PHP 1,200,000 - 1,500,000
Performance bonus (STAR)
Flexible work schedule with work-from-
Health insurance
+2
Site Reliability Engineer
Site Reliability Engineer

PeoplePlusTech Inc. • Metro Manila

Hybrid
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc. • Taguig

On-site
PHP 1,000,000 - 1,500,000