Lead Site Reliability Engineer

iScale Solutions, Inc.

Metro Manila

On-site

PHP 900,000 - 1,800,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary package
Health coverage
Learning opportunities
Team engagement

Job summary

iScale Solutions seeks a results-driven Site Reliability Engineer to lead and manage the software quality and reliability function. You will drive incident response, observability, and scalable architectures for cloud and on‑prem environments.

You will implement SLOs/SLIs, manage IaC with Terraform/Pulumi, and collaborate across Dev and Ops to ensure high‑performance services. Expect a dynamic, global team and continuous improvement.

Qualifications

  • 3–5 years of SRE experience in cloud and on‑prem environments.
  • Strong Linux, networking, and systems administration knowledge.
  • Hands‑on AWS experience with Kubernetes and container orchestration.

Responsibilities

  • Implement and manage SLOs, SLIs, and error budgets to drive reliability.
  • Develop resilient systems ensuring 99.9%+ uptime for critical services.
  • Lead incident response and blameless postmortems with root cause analysis.
  • Automate detection/response with runbooks or workflows.
  • Write production level code in Python or Go as needed.
  • Design observability using OpenTelemetry for traces, metrics, logs.
  • Plan capacity, perform performance testing, ensure scalable growth.
  • Collaborate with development and operations for reliable services.
  • Champion Infrastructure as Code with Terraform/Pulumi.
  • Participate in on‑call rotations.

Skills

SRE experience
Linux
Networking
Systems administration
AWS
Kubernetes
OpenTelemetry
Python
Go
Shell scripting
Observability tools
Terraform
Pulumi
Chaos Engineering

Tools

Terraform
Pulumi

Job description

About the role

You will lead and manage the Software quality assurance function and help maintain and improve the quality of software products, development practices and client implementations. You will be responsible to ensure products are performing and scalable. This means you will create well-structured test plans and test cases and lead the testing activities for your team.

Key responsibilities
  • Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
  • Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
  • Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
  • Automate incident detection and response using automated runbooks or predefined workflows.
  • Write software as needed to support reliability or efficiency needs.
  • Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
  • Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.
  • Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
  • Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
  • Participate on on-call rotation.
About you
  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
  • Deep understanding of Linux systems, networking, and systems administration.
  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
  • Strong skills in at least one programming language (Python, Go) to write production level code.
  • Strong skills in shell scripting using bash or similar.
  • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc).
Benefits
  • Competitive Salary Package: Receive a pay package that matches your skills and experience.
  • Vacation and Sick Leave credits: Enjoy vacation and sick leave credits to maintain work-life balance.
  • Health Coverage: Get medical, dental, and vision insurance for you and your dependents.
  • Government-Mandated Benefits: Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.
  • Learning Opportunities: Access training, certifications, and mentorship to grow your career.
  • Team Engagement: Join team-building activities and wellness programs.
  • Modern Tools: Use the latest technology to excel in your role.
  • Career Growth: Clear paths for promotion and professional development.
  • Inclusive Culture: Be part of a diverse, supportive, and collaborative global team.
  • Referral Rewards: Earn bonuses for bringing great talent to the team.
About us

iScale Solutions is a Managed Outsourcing and Staff Augmentation provider with operations in the Philippines, Madagascar and Singapore. We thrive to provide customized solutions and deeply integrate in our customers' business processes. We believe in recruiting the best available talent in the market, and offer our customers great value for their money.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

iScale Solutions • Philippines

Remote
PHP 7,538,000 - 11,307,000
Competitive salary
Health coverage
Vacation & sick leave
+7
Site Reliability Engineer
Site Reliability Engineer

MicroSourcing • Manila

On-site
PHP 900,000 - 1,500,000
Healthcare coverage
Paid time-off with cash conversion
Group life insurance
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Blackfort Consulting, Inc.. • Pateros

On-site
PHP 1,200,000 - 2,000,000
SRE Lead: Reliability, Observability & Scalable Systems
SRE Lead: Reliability, Observability & Scalable Systems

iScale Solutions, Inc. • Metro Manila

On-site
PHP 900,000 - 1,800,000
Competitive salary package
Health coverage
Learning opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

Coretex • Manila

On-site
PHP 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila, Hinoba-an

On-site
PHP 1,800,000 - 3,000,000
Site Reliability Engineer - LInE (Remote)
Site Reliability Engineer - LInE (Remote)

Quik Hire Staffing • Philippines

Remote
PHP 5,653,000 - 9,422,000
Remote Lead SRE: Scale, Reliability & Observability
Remote Lead SRE: Scale, Reliability & Observability

iScale Solutions • Philippines

Remote
PHP 7,538,000 - 11,307,000
Competitive salary
Health coverage
Vacation & sick leave
+7
Remote Site Reliability Engineer – Scale Cloud Infra
Remote Site Reliability Engineer – Scale Cloud Infra

Philippines (PHC1) Avid Philippines • Philippines

Remote
PHP 1,200,000 - 1,800,000
Health & life insurance
Referral rewards
Generous leave policies
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Gratitude Philippines • Quezon City

Hybrid
PHP 1,000,000 - 1,800,000