Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc.

Taguig

On-site

PHP 1,000,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AIPI Acquire Intelligence Philippines Inc. is seeking a Site Reliability Engineer to ensure the reliability and performance of our IoT telemetry platform. Your responsibilities will include defining Service Level Objectives (SLOs), implementing automated operational processes, and maintaining incident response procedures.

The ideal candidate should have experience with AWS services and monitoring tools, while also ensuring security compliance. This exciting role offers global support coverage and a progressive work environment.

Qualifications

  • Experience managing AWS services, focusing on EKS and MSK.
  • Strong understanding of monitoring tools like Prometheus and Grafana.
  • Proficient in Infrastructure as Code using Pulumi.

Responsibilities

  • Define and enforce Service Level Objectives (SLOs) across systems.
  • Automate operational processes to reduce engineering toil.
  • Serve as incident commander during outages.

Skills

Monitoring and alerting strategies
Infrastructure as Code (IaC) with Pulumi
AWS services management
Incident management
Security best practices

Tools

Prometheus
Grafana
PagerDuty
TypeScript
AWS EKS
MongoDB

Job description

Role Objective

We’re an award‑winning global outsourcer providing contact center and back office services on behalf of our global clients. Come work at a place where innovation and teamwork come together to support the most exciting missions in the world! The Site Reliability Engineer serves as the guardian of our production systems, ensuring the reliability, scalability, and performance of our IoT telemetry platform. You will define and enforce Service Level Objectives (SLOs), automate operational processes, and build the infrastructure and tooling that enables our engineering teams to deploy with confidence. By implementing comprehensive monitoring, incident response procedures, and reliability practices, you will play a pivotal role in maintaining the uptime and data freshness that our customers depend on for their critical fleet operations.

Key Focus Areas
  • SLO Management
  • Infrastructure Automation
  • Incident Response
  • Security & Compliance
Responsibilities
  • Define, monitor, and enforce Service Level Objectives (SLOs) and error budgets across all production systems
  • Track error budget burn rates and make data‑driven decisions to halt risky deployments when thresholds are exceeded
  • Implement comprehensive monitoring and alerting strategies using Prometheus, Grafana, and PagerDuty
  • Establish and maintain reliability standards that support business‑critical uptime requirements
  • Design and implement Infrastructure as Code (IaC) solutions using Pulumi with TypeScript
  • Manage and optimize AWS services including EKS (Elastic Kubernetes Service), MSK (Managed Streaming for Kafka), SingleStore, MongoDB, and S3
  • Automate operational processes to eliminate toil, targeting any task that consumes more than 2 engineer‑days per quarter
  • Serve as incident commander during production outages and service degradations
  • Lead comprehensive post‑mortem processes within 48 hours of incidents
  • Drive "never‑again" corrective actions to completion, ensuring systemic improvements
  • Maintain and improve incident response procedures and runbooks
  • Implement and enforce least‑privilege IAM policies across all AWS resources
  • Manage security patch pipelines and vulnerability remediation processes
  • Support compliance initiatives including SOC2 and ISO 27001 certification requirements
  • Ensure security best practices are embedded in all infrastructure and operational procedures
  • Participate in follow‑the‑sun on‑call rotation with one week primary/secondary commitment every five weeks
  • Provide 24×7 support coverage across AU/NZ, EU/ZA, and MX time zones
  • Maintain operational runbooks and knowledge transfer documentation
  • Continuously improve on‑call experience and reduce alert fatigue
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

LTM • Mexico

On-site
PHP 900,000 - 1,300,000
Site Reliability Engineer
Site Reliability Engineer

Private Advertiser • Makati

On-site
PHP 1,200,000 - 1,800,000
Site Reliability Engineering Onsite
Site Reliability Engineering Onsite

Gratitude Philippines • Manila

On-site
PHP 900,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

NCS Philippines • Taguig

Hybrid
PHP 900,000 - 1,300,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Engineer
Site Reliability Engineer

Alsons/AWS Information Systems Inc. • Cebu City

Hybrid
PHP 600,000 - 1,000,000
IoT Telemetry SRE: Reliability, Automation & 24x7 Ops
IoT Telemetry SRE: Reliability, Automation & 24x7 Ops

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000