Senior Site Reliability Engineer

Maya

Mandaluyong

On-site

PHP 1,800,000 - 2,400,000

Full time

7 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Maya in the Philippines is seeking a seasoned Site Reliability Engineer to lead architectural design and implementation of fault-tolerant, self-healing infrastructure across cloud and hybrid environments. You will own automation strategies, IaC and CI/CD frameworks, and guide reliability initiatives across multiple teams.

You will establish cloud governance, manage budgets, ensure compliance with CIS/PCI-DSS frameworks, and drive cost optimization through FinOps and policy-as-code.

Qualifications

  • Expert-level Kubernetes (CRDs/Operators, multi-tenancy).
  • Advanced Terraform with modules and automated testing.
  • Deep Service Mesh knowledge (Istio).
  • Proven IDP experience with self-service workflows.
  • Advanced GitLab CI/CD and GitOps (ArgoCD/FluxCD).
  • Expert security for APIs (WAF, API Gateway).
  • Strong Go/Python/Java coding and review for reliability.
  • Led cross-functional reliability programs across teams.
  • Proficient observability with Dynatrace/Prometheus/OpenTelemetry.
  • Architected microservices with high availability patterns.
  • Experience with AWS Organizations and multi-account governance.
  • Policy-as-code and compliance automation.
  • Knowledge of CIS, Well-Architected, security standards.
  • Experience with Dynatrace/Datadog/Grafana for enterprise observability.
  • SLO-based alerting and burn-rate monitoring.
  • Distributed tracing and log aggregation (Jaeger/Loki).

Responsibilities

  • Lead architectural decisions for fault-tolerant, self-healing infra.
  • Drive IaC and CI/CD initiatives to automate operations.
  • Own reliability programs across multiple teams and services.
  • Establish cloud governance, cost optimization, and compliance automation.

Skills

Kubernetes expertise
Terraform expertise
Service Mesh (Istio)
Internal Developer Platform (IDP)
GitLab CI/CD / GitOps (ArgoCD/FluxCD)
API gateway security
Go/Python/Java
Cross-functional leadership
Observability platforms (Dynatrace/Dat
Microservices high availability
AWS Organizations / multi-account
Policy-as-code (AWS Config/OPA)
Cloud security standards (CIS, Well-Ar
Dynatrace/Datadog/Grafana
SLO-based alerting & burn rate
Distributed tracing (Jaeger/OpenTeleg)
Log aggregation (ELK/Loki)
Custom metrics design

Tools

Kong
Apigee
AWS API Gateway
Prometheus
Grafana
Jaeger

Job description


  • Lead architectural design and implementation of fault-tolerant, self-healing infrastructure across cloud and hybrid environments

  • Drive organization-wide automation initiatives, eliminating manual operations through advanced IaC and CI/CD frameworks

  • Own technical program leadership for reliability initiatives spanning multiple teams and services

  • Strategic management of OPEX and CAPEX budgets with cost optimization accountability

  • Deep expertise in compliance frameworks (CIS, PCI-DSS, BSP) with ability to architect compliant solutions

  • Establish and enforce cloud governance policies, account structures, and organizational standards across AWS/Azure/GCP environments


DISPLAYED SKILL MASTERY


  • Architect and implement advanced CI/CD pipelines with progressive delivery patterns (canary, blue-green)

  • Design and maintain enterprise-grade Infrastructure-as-Code modules with reusability and governance

  • Lead complex incident resolution and conduct deep-dive root cause analysis

  • Drive adoption of emerging technologies and reliability patterns across engineering teams

  • Mentor Senior SREs on architectural decisions and reliability best practices

  • Design and implement cloud landing zones, multi-account strategies, and policy-as-code frameworks

  • Build comprehensive SLI/SLO frameworks with automated alerting, error budget tracking, and burn rate analysis

  • Correlate metrics across distributed systems using APM, distributed tracing, and custom dashboards


EXPECTED RESULTS


  • Build and maintain self-service platforms enabling developer autonomy while ensuring reliability standards

  • Implement chaos engineering practices and resilience testing frameworks

  • Lead R&D initiatives evaluating emerging technologies for production adoption

  • Define and enforce SLI/SLO frameworks across services, driving error budget policies

  • Architect solutions that materially improve system availability, performance, and cost efficiency

  • Conduct capacity planning and optimization initiatives across infrastructure domains

  • Influence engineering decisions through architectural reviews and reliability assessments

  • Enforce cloud governance through SCPs, IAM policies, resource tagging standards, and compliance automation

  • Drive cloud cost optimization through FinOps practices, rightsizing, and resource lifecycle management

  • Establish service health scoring systems with predictive alerting and anomaly detection

  • Create self-service observability tooling enabling teams to define and monitor their own SLOs


REQUIRED QUALIFICATIONS


  • Expert-level proficiency in Kubernetes (CRDs, Operators, multi-tenancy, advanced scheduling)

  • Advanced Terraform expertise (custom providers, module design, automated testing)

  • Deep Service Mesh knowledge (Istio traffic management, circuit breaking, rate limiting, mTLS)

  • Proven experience building Internal Developer Platforms (IDP) with self-service workflows

  • Advanced GitLab CI/CD and GitOps implementation (ArgoCD/FluxCD, multi-project pipelines)

  • Expert-level WAF, API Gateway (Kong, Apigee, AWS APIGW), and network security implementation

  • Strong software development skills in Go, Python, or Java with ability to review code for reliability impact

  • Experience leading technical programs and cross-functional reliability initiatives

  • Deep understanding of observability platforms (Dynatrace, Prometheus, OpenTelemetry) with custom integration experience

  • Proven track record architecting microservices with high-availability and resiliency patterns

  • Experience implementing AWS Organizations, Control Tower, Service Control Policies, and multi-account governance frameworks

  • Proficiency in cloud policy-as-code tools (AWS Config, OPA, Sentinel) and compliance automation

  • Knowledge of cloud security standards (CIS Benchmarks, AWS Well-Architected Framework, Azure/GCP best practices)

  • Advanced expertise in Dynatrace, Datadog, or Grafana for building enterprise observability solutions

  • Experience implementing SLO-based alerting, error budgets, and burn rate monitoring using Prometheus, Grafana, or commercial APM tools

  • Proficiency in distributed tracing (Jaeger, Zipkin, OpenTelemetry) and log aggregation (ELK, Loki)

  • Ability to design custom metrics, synthetic monitoring, and real user monitoring (RUM) strategies

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer- AI-Driven SRE & Cloud SRE
Software Engineer- AI-Driven SRE & Cloud SRE

Keka Technologies Private Limited • Mexico

On-site
PHP 2,215,000 - 3,322,000
Senior Devops Engineer
Senior Devops Engineer

V2 Solutions • Hinoba-an

Hybrid
PHP 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

PeoplePlusTech Inc. • Metro Manila

Hybrid
PHP 700,000 - 1,100,000
Senior Site Reliability Engineer SRE Kubernetes
Senior Site Reliability Engineer SRE Kubernetes

Accenture in the Philippines • Cebu City

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000
SRE Manager
SRE Manager

Quality Ai • Hinoba-an

On-site
PHP 2,653,000 - 4,642,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA • Philippines

On-site
PHP 900,000 - 1,500,000
Application Support Engineer
Application Support Engineer

Accenture in the Philippines • Taguig

On-site
PHP 1,200,000 - 2,200,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Accenture • Cebu City

On-site
PHP 700,000 - 1,100,000