DevOps & SRE Engineer—Cloud Reliability & Observability

iScale Solutions

Philippines

On-site

PHP 900,000 - 1,500,000

Full time

15 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

iScale Solutions seeks an experienced SRE to design and operate reliable, scalable services in cloud and on-prem environments. You will own SLOs/SLIs, lead incident response, and drive observability using OpenTelemetry, Grafana, and Prometheus.

The role requires programming skills (Python/Go), IaC (Terraform/Pulumi), and chaos engineering practice. Collaborating with development and operations teams, you’ll ensure robust uptime and efficient deployments while contributing to capacity planning

Qualifications

  • 3-5 years as an SRE in cloud and on-prem environments.
  • Strong Linux, networking and systems administration skills.
  • Experience with AWS and Kubernetes; proficient in observability tooling.
  • Proficient in at least one programming language (Python or Go).
  • Shell scripting proficiency and knowledge of distributed tracing.
  • Familiarity with Chaos Engineering methods and tools.
  • Experience with IaC and on-call incident management.

Responsibilities

  • Implement and manage SLOs/SLIs and error budgets for reliability.
  • Develop systems for 99.9%+ uptime for critical services.
  • Lead incident response and postmortems with root cause analysis.
  • Automate incident detection and runbooks; write supporting software as needed.
  • Design full observability using OpenTelemetry for tracing, metrics, and logs.
  • Plan capacity, forecast demand, and test performance for scale.
  • Collaborate across teams to build reliable, scalable services.
  • Apply best practices for IaC in provisioning and managing infrastructure.
  • Engage in chaos engineering initiatives and participate in on-call rotations.
  • Drive advanced alerting and anomaly detection on metrics.

Skills

SRE experience
Linux administration
Networking
Shell scripting
Incident response
On-call experience
Capacity planning
Python/Go programming

Tools

AWS
Kubernetes
OpenTelemetry
Grafana
Prometheus
Thanos
ELK / Elastic
Loki
Terraform
Pulumi
Chaos Mesh
Chaos Monkey
AWS Fault Injection

Job description

iScale Solutions seeks an experienced SRE to design and operate reliable, scalable services in cloud and on-prem environments. You will own SLOs/SLIs, lead incident response, and drive observability using OpenTelemetry, Grafana, and Prometheus.

The role requires programming skills (Python/Go), IaC (Terraform/Pulumi), and chaos engineering practice. Collaborating with development and operations teams, you’ll ensure robust uptime and efficient deployments while contributing to capacity planning

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Lead: Reliability, Observability & Scalable Systems
SRE Lead: Reliability, Observability & Scalable Systems

iScale Solutions, Inc. • Metro Manila

On-site
PHP 900,000 - 1,800,000
Competitive salary package
Health coverage
Learning opportunities
+1
Remote Lead SRE: Scale, Reliability & Observability
Remote Lead SRE: Scale, Reliability & Observability

iScale Solutions • Philippines

Remote
PHP 7,538,000 - 11,307,000
Competitive salary
Health coverage
Vacation & sick leave
+7
Lead DevOps Engineer
Lead DevOps Engineer

iScale Solutions • Philippines

On-site
PHP 900,000 - 1,500,000
Reliability Engineer: SRE, Observability & Cloud Ops
Reliability Engineer: SRE, Observability & Cloud Ops

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000
Lead DevOps Engineer
Lead DevOps Engineer

TymblHub • Hinoba-an

On-site
PHP 1,004,000 - 1,451,000
L1 SRE Engineer - Cloud & DevOps Operations
L1 SRE Engineer - Cloud & DevOps Operations

TymblHub • Hinoba-an

On-site
PHP 900,000 - 1,300,000
Global SRE — Scale Observability & Automation
Global SRE — Scale Observability & Automation

Vestas • Philippines

On-site
PHP 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

TymblHub • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Senior Devops Engineer
Senior Devops Engineer

TymblHub • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000
Senior DevOps & SRE | Hybrid Kubernetes Architect
Senior DevOps & SRE | Hybrid Kubernetes Architect

TymblHub • Hinoba-an

Hybrid
PHP 1,200,000 - 2,400,000