Lead Site Reliability Engineer

iScale Solutions

Philippines

Remote

PHP 7,538,000 - 11,307,000

Full time

12 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Health coverage
Vacation & sick leave
Government-mandated benefits
Training & certifications
Team engagement
Modern tools
Career growth
Inclusive culture
Referral rewards

Job summary

iScale Solutions is seeking an experienced Site Reliability Engineer for a remote role. You will build and operate scalable, reliable services, focusing on SLOs, SLI, and incident response across cloud and on‑prem environments.

You will implement full observability using OpenTelemetry, manage automation, and participate in on-call rotations. Strong Linux, AWS, Kubernetes, and programming skills (Python/Go) are required, along with chaos engineering experience.

Qualifications

  • 3-5 years of experience as an SRE in cloud and on‑prem environments.
  • Deep understanding of Linux systems, networking, and sysadmin.
  • Hands-on with AWS and Kubernetes/container orchestration.
  • Experience with observability tools (Grafana, Prometheus, Honeycomb, ELK, Loki).
  • Proficient in Python or Go for production code.
  • Shell scripting (bash) and OpenTelemetry or distributed tracing.
  • Chaos engineering experience (Chaos Mesh, AWS FIS, chaos monkey).

Responsibilities

  • Implement and manage SLOs, SLIs, and error budgets.
  • Develop resilient systems with 99.9%+ uptime.
  • Lead incident response and blameless postmortems.
  • Automate incident detection and runbooks.
  • Write software to support reliability or efficiency.
  • Design observability across systems using OpenTelemetry.
  • Plan capacity, forecast demand, and perform performance testing.
  • Collaborate with development and operations on reliable services.
  • Follow best practices for infrastructure design, deployment, and maintenance (AWS, Kubernetes, EKS, Fargate).
  • Champion Infrastructure as Code (Terraform, Pulumi) to provision and scale.
  • Participate in chaos engineering initiatives.
  • Join on-call rotation.
  • Drive advanced alerting and anomaly detection on metrics.

Skills

SRE
Linux
Networking
Python
Go
Shell scripting
OpenTelemetry

Tools

Grafana
Prometheus
Honeycomb
ELK
Loki
Thanos
OpenTelemetry
Chaos Mesh
AWS
Kubernetes
Terraform
Pulumi

Job description

This is a remote position.

  • Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.

  • Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.

  • Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.

  • Automate incident detection and response using automated runbooks or predefined workflows.

  • Write software as needed to support reliability or efficiency needs.

  • Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.

  • Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.

  • Collaborate with development and operations teams on building reliable, scalable, and high-performance services.

  • Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.

  • Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.

  • Get involved in chaos engineering initiatives.

  • Participate on our on-call rotation.

  • Drive advanced alerting and anomaly detection applied to metrics

Requirements
  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.

  • Deep understanding of Linux systems, networking, and systems administration.

  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.

  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.

  • Strong skills in at least one programming language (Python, Go) to write production level code.

  • Strong skills in shell scripting using bash or similar. Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.

  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.

Benefits
  • Competitive Salary Package: Receive a pay package that matches your skills and experience.

  • Vacation and Sick Leave credits: Enjoy vacation and sick leave credits to maintain work-life balance.

  • Health Coverage: Get medical, dental, and vision insurance for you and your dependents.

  • Government-Mandated Benefits: Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.

  • Learning Opportunities: Access training, certifications, and mentorship to grow your career.

  • Team Engagement: Join team-building activities and wellness programs.

  • Modern Tools: Use the latest technology to excel in your role.

  • Career Growth: Clear paths for promotion and professional development.

  • Inclusive Culture: Be part of a diverse, supportive, and collaborative global team.

  • Referral Rewards: Earn bonuses for bringing great talent to the team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

iScale Solutions, Inc. • Metro Manila

On-site
PHP 900,000 - 1,800,000
Competitive salary package
Health coverage
Learning opportunities
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Blackfort Consulting, Inc.. • Pateros

On-site
PHP 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

MicroSourcing • Manila

On-site
PHP 900,000 - 1,500,000
Healthcare coverage
Paid time-off with cash conversion
Group life insurance
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Gratitude Philippines • Quezon City

Hybrid
PHP 1,000,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Coretex • Manila

On-site
PHP 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila, Hinoba-an

On-site
PHP 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Philtech Inc. • Taguig

On-site
PHP 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Gratitude Philippines • Quezon City

Hybrid
PHP 1,200,000 - 1,800,000
Hybrid work arrangement
Night shift
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,004 - 7,371,007
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Engineer (SRE) - AWS & Kurbernetes
Site Reliability Engineer (SRE) - AWS & Kurbernetes

Lewis Personnel Management • Metro Manila

Remote
PHP 900,000 - 1,500,000