Lead DevOps Engineer

iScale Solutions

Philippines

On-site

PHP 900,000 - 1,500,000

Full time

23 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

iScale Solutions seeks an experienced SRE to design and operate reliable, scalable services in cloud and on-prem environments. You will own SLOs/SLIs, lead incident response, and drive observability using OpenTelemetry, Grafana, and Prometheus.

The role requires programming skills (Python/Go), IaC (Terraform/Pulumi), and chaos engineering practice. Collaborating with development and operations teams, you’ll ensure robust uptime and efficient deployments while contributing to capacity planning

Qualifications

  • 3-5 years as an SRE in cloud and on-prem environments.
  • Strong Linux, networking and systems administration skills.
  • Experience with AWS and Kubernetes; proficient in observability tooling.
  • Proficient in at least one programming language (Python or Go).
  • Shell scripting proficiency and knowledge of distributed tracing.
  • Familiarity with Chaos Engineering methods and tools.
  • Experience with IaC and on-call incident management.

Responsibilities

  • Implement and manage SLOs/SLIs and error budgets for reliability.
  • Develop systems for 99.9%+ uptime for critical services.
  • Lead incident response and postmortems with root cause analysis.
  • Automate incident detection and runbooks; write supporting software as needed.
  • Design full observability using OpenTelemetry for tracing, metrics, and logs.
  • Plan capacity, forecast demand, and test performance for scale.
  • Collaborate across teams to build reliable, scalable services.
  • Apply best practices for IaC in provisioning and managing infrastructure.
  • Engage in chaos engineering initiatives and participate in on-call rotations.
  • Drive advanced alerting and anomaly detection on metrics.

Skills

SRE experience
Linux administration
Networking
Shell scripting
Incident response
On-call experience
Capacity planning
Python/Go programming

Tools

AWS
Kubernetes
OpenTelemetry
Grafana
Prometheus
Thanos
ELK / Elastic
Loki
Terraform
Pulumi
Chaos Mesh
Chaos Monkey
AWS Fault Injection

Job description

  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
  • Deep understanding of Linux systems, networking, and systems administration.
  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
  • Strong skills in at least one programming language (Python, Go) to write production level code.
  • Strong skills in shell scripting using bash or similar.
  • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault
Responsibilities:
  • Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
  • Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
  • Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
  • Automate incident detection and response using automated runbooks or predefined workflows.
  • Write software as needed to support reliability or efficiency needs.
  • Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
  • Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.
  • Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
  • Ensure best practices are followed across infrastructure design, deployment, and maintenance using
  • Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
  • Get involved in chaos engineering initiatives.
  • Participate on our on‑call rotation.
  • Drive advanced alerting and anomaly detection applied to metrics
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead DevOps Engineer
Lead DevOps Engineer

TymblHub • Hinoba-an

On-site
PHP 1,004,000 - 1,451,000
Senior Devops Engineer
Senior Devops Engineer

TymblHub • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000
DevOps & SRE Engineer—Cloud Reliability & Observability
DevOps & SRE Engineer—Cloud Reliability & Observability

iScale Solutions • Philippines

On-site
PHP 900,000 - 1,500,000
Application Support Engineer
Application Support Engineer

Accenture in the Philippines • Taguig

On-site
PHP 1,200,000 - 2,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Blackfort Consulting, Inc.. • Pateros

On-site
PHP 1,200,000 - 2,000,000
DevOps Engineer Lead _100% Remote_Nightshift _Up to 300K
DevOps Engineer Lead _100% Remote_Nightshift _Up to 300K

weSource Management Consultancy Firm • Metro Manila

Remote
PHP 270,000 - 330,000
DevOps Lead
DevOps Lead

NCS Group • Taguig

On-site
PHP 1,800,000 - 3,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

iScale Solutions • Philippines

Remote
PHP 7,538,000 - 11,307,000
Competitive salary
Health coverage
Vacation & sick leave
+7
Senior DevOps Engineer
Senior DevOps Engineer

Cloud Bridge • Philippines

On-site
PHP 1,200,000 - 2,400,000