Senior Site Reliability Engineer

Hammerjack Pty Ltd

Philippines

On-site

PHP 1,800,000 - 2,400,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hammerjack Pty Ltd in the Philippines seeks a senior SRE/Platform Engineer to lead fault-tolerant, self-healing infrastructure across cloud and hybrid environments.

You will drive IaC and CI/CD automation, own reliability programs across multiple teams, and manage budgets with cost optimization accountability, while aligning with CIS/PCI-DSS and BSP compliance. Establish cloud governance across AWS/Azure/GCP and build an internal developer platform with self-service workflows.

Qualifications

  • 7-10+ years experience focused on SRE, DevOps, Platform Engineering, or Cloud Infrastructure
  • Expert-level proficiency in Kubernetes (CRDs, Operators, multi-tenancy, advanced scheduling)
  • Advanced Terraform expertise (custom providers, module design, automated testing)
  • Deep Service Mesh knowledge (Istio traffic management, circuit breaking, rate limiting, mTLS)
  • Proven experience building Internal Developer Platforms (IDP) with self-service workflows
  • Advanced GitLab CI/CD and GitOps implementation (ArgoCD/FluxCD, multi-project pipelines)
  • Expert-level WAF, API Gateway (Kong, Apigee, AWS APIGW), and network security implementation
  • Strong software development skills in Go, Python, or Java with ability to review code for reliability impact
  • Experience leading technical programs and cross-functional reliability initiatives
  • Deep understanding of observability platforms (Dynatrace, Prometheus, OpenTelemetry) with custom integration experience
  • Proven track record architecting microservices with high-availability and resiliency patterns
  • Experience implementing AWS Organizations, Control Tower, Service Control Policies, and multi-account governance frameworks
  • Proficiency in cloud policy-as-code tools (AWS Config, OPA, Sentinel) and compliance automation
  • Knowledge of cloud security standards (CIS Benchmarks, AWS Well-Architected Framework, Azure/GCP best practices)
  • Advanced expertise in Dynatrace, Datadog, or Grafana for building enterprise observability solutions
  • Experience implementing SLO-based alerting, error budgets, and burn rate monitoring using Prometheus, Grafana, or commercial APM tools
  • Proficiency in distributed tracing (Jaeger, Zipkin, OpenTelemetry) and log aggregation (ELK, Loki)
  • Ability to design custom metrics, synthetic monitoring, and real user monitoring (RUM) strategies

Responsibilities

  • Lead architectural design and implementation of fault-tolerant, self-healing infrastructure across cloud and hybrid environments
  • Drive organization-wide automation initiatives with IaC and CI/CD frameworks
  • Own technical program leadership for reliability initiatives spanning multiple teams and services
  • Strategic management of OPEX and CAPEX budgets with cost optimization accountability
  • Establish and enforce cloud governance policies, account structures, and organizational standards across AWS/Azure/GCP environments

Skills

Kubernetes
Terraform
Istio
GitLab CI/CD
Go/Python/Java
Observability
SRE/Platform Engineering
Cloud Governance

Tools

ArgoCD/FluxCD
Dynatrace/Grafana/Datadog
Prometheus
Jaeger/Zipkin/OpenTelemetry
Kong/Apigee/AWS API Gateway

Job description

NATURE OF WORK

  • Lead architectural design and implementation of fault-tolerant, self-healing infrastructure across cloud and hybrid environments
  • Drive organization-wide automation initiatives, eliminating manual operations through advanced IaC and CI/CD frameworks
  • Own technical program leadership for reliability initiatives spanning multiple teams and services
  • Strategic management of OPEX and CAPEX budgets with cost optimization accountability
  • Deep expertise in compliance frameworks (CIS, PCI-DSS, BSP) with ability to architect compliant solutions
  • Establish and enforce cloud governance policies, account structures, and organizational standards across AWS/Azure/GCP environments

REQUIRED QUALIFICATIONS

  • 7-10+ years experience focused on SRE, DevOps, Platform Engineering, or Cloud Infrastructure
  • Expert-level proficiency in Kubernetes (CRDs, Operators, multi-tenancy, advanced scheduling)
  • Advanced Terraform expertise (custom providers, module design, automated testing)
  • Deep Service Mesh knowledge (Istio traffic management, circuit breaking, rate limiting, mTLS)
  • Proven experience building Internal Developer Platforms (IDP) with self-service workflows
  • Advanced GitLab CI/CD and GitOps implementation (ArgoCD/FluxCD, multi-project pipelines)
  • Expert-level WAF, API Gateway (Kong, Apigee, AWS APIGW), and network security implementation
  • Strong software development skills in Go, Python, or Java with ability to review code for reliability impact
  • Experience leading technical programs and cross-functional reliability initiatives
  • Deep understanding of observability platforms (Dynatrace, Prometheus, OpenTelemetry) with custom integration experience
  • Proven track record architecting microservices with high-availability and resiliency patterns
  • Experience implementing AWS Organizations, Control Tower, Service Control Policies, and multi-account governance frameworks
  • Proficiency in cloud policy-as-code tools (AWS Config, OPA, Sentinel) and compliance automation
  • Knowledge of cloud security standards (CIS Benchmarks, AWS Well-Architected Framework, Azure/GCP best practices)
  • Advanced expertise in Dynatrace, Datadog, or Grafana for building enterprise observability solutions
  • Experience implementing SLO-based alerting, error budgets, and burn rate monitoring using Prometheus, Grafana, or commercial APM tools
  • Proficiency in distributed tracing (Jaeger, Zipkin, OpenTelemetry) and log aggregation (ELK, Loki)
  • Ability to design custom metrics, synthetic monitoring, and real user monitoring (RUM) strategies
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Maya • Mandaluyong

On-site
PHP 1,800,000 - 2,400,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000
Senior Engineer - Site Reliability
Senior Engineer - Site Reliability

Dencom Consultancy and Manpower Services • Parañaque

On-site
Senior Devops Engineer
Senior Devops Engineer

V2 Solutions • Hinoba-an

Hybrid
PHP 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA • Philippines

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

V2 Solutions • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Accenture • Cebu City

On-site
PHP 700,000 - 1,100,000
Staff Site Reliability Engineer – Cloud Efficiency
Staff Site Reliability Engineer – Cloud Efficiency

Super • España

On-site
PHP 1,200,000 - 1,600,000
Medical / Health Insurance
Employee Assistance Programme
Site Reliability Engineer
Site Reliability Engineer

Alsons/AWS Information Systems Inc. • Cebu City

Hybrid
PHP 600,000 - 1,000,000