Senior SRE – Cloud SaaS & Observability

Cloud Blue

Philippines

On-site

PHP 5,058,000 - 7,948,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work
Competitive salary
Career development
Flexible hours

Job summary

HostPapa's CloudBlue team seeks a Site Reliability Engineer to bolster reliability, scalability, and observability of multi-tenant SaaS platforms used by service providers worldwide. You will define SLIs/SLOs, build a robust observability stack with Datadog, Grafana, and Elasticsearch/Kibana, and lead incident response across Kubernetes-based systems.

A strong background in Linux, Docker, Kubernetes, Python or Bash, and on-call experience will help you succeed in this remote role with

Qualifications

  • 3+ years as an SRE, DevOps Engineer, or Production Engineer
  • Experience operating highly available, multi-tenant SaaS platforms
  • Hands-on with observability/monitoring tools such as Datadog, Grafana, Elasticsearch/Kibana
  • Solid Linux, networking, and distributed systems fundamentals
  • Experience with Docker and Kubernetes
  • Strong scripting/automation skills (Python/Bash)
  • On-call rotations and incident response in production environments
  • Fluent written/spoken English

Responsibilities

  • Define and implement SLIs, SLOs, and error budgets for CloudBlue services
  • Influence architecture for reliability, scalability, and operability
  • Reduce toil with automation and process improvements
  • Design and operate observability stack across metrics, logs, traces
  • Develop alerting strategies and dashboards for health insights
  • Maintain high-availability architectures across regions
  • Capacity planning, load testing, and performance optimization
  • Lead blameless postmortems and drive improvements
  • Improve reliability of Kubernetes platforms with health checks and autoscaling
  • Collaborate to improve deployment safety and rollback strategies
  • Maintain runbooks and promote SRE best practices
  • Support additional tasks to meet team needs

Skills

SRE ownership
Observability
Linux fundamentals
Python/Bash
Incident response
On-call experience
English communication

Tools

Datadog
Grafana
Elasticsearch/Kibana
Docker
Kubernetes

Job description

HostPapa's CloudBlue team seeks a Site Reliability Engineer to bolster reliability, scalability, and observability of multi-tenant SaaS platforms used by service providers worldwide. You will define SLIs/SLOs, build a robust observability stack with Datadog, Grafana, and Elasticsearch/Kibana, and lead incident response across Kubernetes-based systems.

A strong background in Linux, Docker, Kubernetes, Python or Bash, and on-call experience will help you succeed in this remote role with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud Reliability, Automation & Observability
Senior SRE: Cloud Reliability, Automation & Observability

Broadridge Financial Solutions • Manila, Hinoba-an

On-site
PHP 1,800,000 - 2,600,000
Senior Cloud SRE — AWS, Automation & Observability
Senior Cloud SRE — AWS, Automation & Observability

Broadridge • Metro Manila

On-site
PHP 1,200,000 - 2,400,000
Cloud Platform & SRE Engineer — Kubernetes & Automation
Cloud Platform & SRE Engineer — Kubernetes & Automation

Global Recruitment and Consultancy OPC • Cebu City

On-site
PHP 1,200,000 - 2,400,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Senior SRE - Kubernetes & Cloud Reliability (Mexico)
Senior SRE - Kubernetes & Cloud Reliability (Mexico)

Sezzle • Mexico

Hybrid
Senior SRE Operations Engineer - Cloud & Observability
Senior SRE Operations Engineer - Cloud & Observability

Teciem • Hinoba-an

Hybrid
PHP 2,772,000 - 4,291,000
Senior SRE: Scale, Automate & Self-Healing Systems
Senior SRE: Scale, Automate & Self-Healing Systems

Replit • España

On-site
PHP 2,461,000 - 4,308,000
Competitive Salary & Equity
Health, Dental, Vision and Life Insurance
Flexible Time Off (FTO) + Holidays
+2
Hybrid Cloud SRE: Observability & Automation
Hybrid Cloud SRE: Observability & Automation

NICE • Manila

Hybrid
PHP 669,600 - 892,800
Reliability Engineer: SRE, Observability & Cloud Ops
Reliability Engineer: SRE, Observability & Cloud Ops

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000