Site Reliability Engineer (SRE)

Acquire Intelligence

Taguig

On-site

PHP 900,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Acquire Intelligence is seeking a Site Reliability Engineer to safeguard our IoT telemetry platform and ensure reliability, scalability, and performance of production systems.

You will define and enforce SLOs, automate operational processes, and build tooling to enable confident deployments. The role emphasizes monitoring, incident response, and security as core pillars of uptime and data freshness for fleet operations.

Qualifications

  • Design and implement IaC solutions (Pulumi/TypeScript) to automate infrastructure.
  • Define and enforce SLOs with error budgets across production systems.
  • Lead incident response and post-mortems with data-driven actions.
  • Maintain security patches, IAM least-privilege policies, and compliance readiness (SOC2, ISO27001).

Responsibilities

  • Define, monitor, and enforce SLOs and error budgets across all production systems.
  • Track error budget burn rates and halt risky deployments when thresholds are exceeded.
  • Implement monitoring and alerting with Prometheus, Grafana, and PagerDuty.
  • Establish reliability standards for business-critical uptime.

Skills

SRE fundamentals
Infrastructure as Code
Pulumi
TypeScript
Monitoring
Incident management
Security best practices
SOC2 / ISO27001

Tools

Prometheus
Grafana
PagerDuty
Pulumi
AWS

Job description

We’re an award-winning global outsourcer providing contact center and back office services on behalf of our global clients. Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!

Role objective

The Site Reliability Engineer serves as the guardian of our production systems, ensuring the reliability, scalability, and performance of our IoT telemetry platform. You will define and enforce Service Level Objectives (SLOs), automate operational processes, and build the infrastructure and tooling that enables our engineering teams to deploy with confidence. By implementing comprehensive monitoring, incident response procedures, and reliability practices, you will play a pivotal role in maintaining the uptime and data freshness that our customers depend on for their critical fleet operations.

The Role Will Focus On The Following Key Areas
  • SLO Management
  • Infrastructure Automation
  • Incident Response
  • Security & compliance
Key Responsibilities

Responsibilities of the Site Reliability Engineer will include but are not limited to:

  • Define, monitor, and enforce Service Level Objectives (SLOs) and error budgets across all production systems
  • Track error budget burn rates and make data-driven decisions to halt risky deployments when thresholds are exceeded
  • Implement comprehensive monitoring and alerting strategies using Prometheus, Grafana, and PagerDuty
  • Establish and maintain reliability standards that support business-critical uptime
Requirements
  • Infrastructure Automation & Management
  • Design and implement Infrastructure as Code (IaC) solutions using Pulumi with TypeScript
  • Manage and optimize AWS services including EKS (Elastic Kubernetes Service), MSK (Managed Streaming for Kafka), SingleStore, MongoDB S3
  • Automate operational processes to eliminate toil, targeting any task that consumes more than 2 engineer-days per quarter
Incident Response & Post-Mortem Leadership
  • Serve as incident commander during production outages and service degradations
  • Lead comprehensive post-mortem processes within 48 hours of incidents
  • Drive "never-again" corrective actions to completion, ensuring systemic improvements
  • Maintain and improve incident response procedures and runbooks
Security & Compliance
  • Implement and enforce least-privilege IAM policies across all AWS resources
  • Manage security patch pipelines and vulnerability remediation processes
  • Support compliance initiatives including SOC2 and ISO 27001 certification requirements
  • Ensure security best practices are embedded in all infrastructure and operational procedures
On-Call & Operational Excellence
  • Participate in follow-the-sun on-call rotation with one week primary/secondary commitment every five weeks
  • Provide 24×7 support coverage across AU/NZ, EU/ZA, and MX time zones
  • Maintain operational runbooks and knowledge transfer documentation
  • Continuously improve on-call experience and reduce alert fatigue

Join the A-Team and experience the A-Life!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc. • Taguig

On-site
PHP 1,000,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

LTM • Mexico

On-site
PHP 900,000 - 1,300,000
Site Reliability Engineer
Site Reliability Engineer

PeoplePlusTech Inc. • Metro Manila

Hybrid
PHP 900,000 - 1,500,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000
Site Reliability Engineering Onsite
Site Reliability Engineering Onsite

Gratitude Philippines • Manila

On-site
PHP 900,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Private Advertiser • Makati

On-site
PHP 1,200,000 - 1,800,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Site Reliability Engineer
Site Reliability Engineer

NCS Philippines • Taguig

Hybrid
PHP 900,000 - 1,300,000
Senior Site Reliability Engineer (AWS)
Senior Site Reliability Engineer (AWS)

broadridge • Philippines

Hybrid
PHP 1,000,000 - 2,400,000
Hybrid work model