Senior Site Reliability Engineer

8x8inc

Manila

On-site

PHP 1,200,000 - 1,600,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

8x8inc is seeking a Senior Reliability Engineer to own platform reliability across global UC infrastructure, lead incident response, and drive automation to reduce toil. You will triage complex issues, participate in post-mortems, and collaborate with diverse teams to improve architecture and tooling.

The role emphasizes on-call leadership, cross-team coordination, and building data-driven reliability practices with SLIs/SLOs, dashboards, and runbooks as core outcomes.

Qualifications

  • 6+ years in site reliability, platform operations, or infrastructure engineering with production-scale systems.
  • Mastery of Linux systems administration: multi-service distributed systems, log reading, systemctl, network diagnostics.
  • Hands-on experience with at least one major cloud provider (OCI/AWS/GCP/Azure).
  • Strong on-call experience: calm under pressure, clear incident communication, lead multi-party responses.

Responsibilities

  • Own platform reliability across global UC infrastructure and drive incident response strategy.
  • Triage and resolve hard issues and act as senior escalation point for NOC and engineers.
  • Lead blameless post-mortems and ensure follow-up actions are completed.
  • Drive automation to reduce toil and improve tooling across subsystems.
  • Define and track SLIs/SLOs; build dashboards (Grafana, OCI Log Analytics).
  • Mentor junior engineers and contribute to knowledge sharing.

Skills

SRE experience
Linux admin
Cloud experience
On-call experience
Scripting (Python/Bash)
Incident response discipline

Tools

PagerDuty
Jira
OCI Log Analytics
Grafana

Job description

8x8 connects our customers and teams globally, empowering CX leaders with performance and insights to make smarter decisions, delight customers, and drive lasting business impact.

What You'll Do
Production Operations & Incident Response

Own platform reliability across global UC infrastructure, driving incident response and the overall reliability strategy for your subsystem rather than resolving issues in isolation.

Triage and resolve the hardest issues - service restarts, hung processes, infrastructure failures - and act as the senior escalation point for the NOC and for other engineers when frontline teams hit their limit.

Execute and improve the unglamorous but essential work: scheduled maintenance, certificate renewals, log rotation - and redesign these processes so failure is prevented systemically, not handled case by case.

Lead blameless post-mortems that produce real follow-through, and sign off on the corrective actions that come out of them.

Cross-Team Collaboration

Work directly with Support, Sales, Sales Engineering, NOC, Professional Services, and Engineering teams across 8x8 - this team sits at the operational center of the company.

Translate production events into clear, business-readable communication under pressure; stakeholders across the org depend on your judgment during incidents.

Feed operational insight back into engineering - turning recurring failures and patterns into actionable bug reports, platform improvements, and influence over the architectural roadmap.

Work closely with technical leads to align reliability and automation work with broader engineering goals, and help focus discussion on what matters most.

Reliability Engineering & Automation

Identify recurring manual work and build automation to eliminate it - we treat toil as a bug, not a requirement.

Drive design for the tooling and automation in your domain; anticipate how a change in one component impacts others and account for adjacent domains in your designs.

Understand the limits of our existing tools - and recognize when a problem exceeds those limits and deserves the effort of building a new one.

Take on large-scale technical debt and refactoring across the subsystem, and contribute to the team's coding methodologies and best practices.

Participate in 2-week sprint cycles to deliver automation, tooling improvements, runbook development, and infrastructure initiatives from a structured backlog. Own the functional specifications for large features and sign off on test plans.

Address security issues as they arise - CVEs, misconfigurations, access control gaps - treated as first-class work alongside incident response.

Define and track SLIs, SLOs, and SLAs to drive honest, data-driven conversations about where reliability investment is needed.

Build and maintain dashboards (Grafana, OCI Log Analytics) that give the team genuine signal; tune alerting to eliminate noise - a high-noise on-call is itself a reliability failure.

Leverage AI-powered tooling to accelerate diagnostics and reduce cognitive load at scale.

Technical Leadership & Mentorship

Provide technical leadership for projects involving 1-2 other engineers.

Consistently mentor more junior engineers; be the person other developers seek out for constructive, insightful feedback.

Frequently and actively share knowledge - of your own work, of areas you've worked in, and of obscure corners outside your immediate context - and encourage others to do the same.

Run workshops, contribute to how-to guides, present at demos, and contribute to the team's presentation portfolio.

On-Call & Coverage

Shared on-call rotation, approximately 1 week per month - same expectation for every engineer on the team.

Escalation is always an option and is encouraged; you are expected to drive the response, set the pace for others, and know when to pull people in - not to hero it alone.

Tooling : PagerDuty for alerting, Jira for tracking, OCI Log Analytics and Grafana for diagnostics.

What We're Looking For
Required
  • 6+ years in a site reliability, platform operations, or infrastructure engineering role - you have run production systems at scale and have a track record of driving, not just maintaining.
  • Mastery of Linux systems administration: multi-service distributed systems, log reading, systemctl, network diagnostics, no GUI required.
  • Deep hands-on experience with at least one major cloud provider (OCI, AWS, GCP, or Azure) - compute, storage, IAM, networking fundamentals.
  • Strong on-call experience: calm under pressure, fast triage, clear communication during an incident, and the judgment to lead a multi‑party response.
  • Scripting in Python or Bash - enough to automate a task, parse logs, hitanAPI, and refactor someoneelse's tooling without breaking it.
  • Strongincidentresponse discipline—structuredthinking…
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8 • Manila, Hinoba-an

On-site
PHP 1,000,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila, Hinoba-an

On-site
PHP 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

8x8, Inc. • Manila

On-site
PHP 781,200 - 1,674,000
Onboarding program
Global-scale production exposure
Blameless post-mortems culture
+1
Senior Site Reliability Engineer - Lead Automation
Senior Site Reliability Engineer - Lead Automation

8x8inc • Manila

On-site
PHP 1,200,000 - 1,600,000
Senior Site Reliability Engineer, Global Infra & Automation
Senior Site Reliability Engineer, Global Infra & Automation

8x8 • Manila, Hinoba-an

On-site
PHP 1,000,000 - 2,000,000
Reliability Operations Engineer (Philippines)
Reliability Operations Engineer (Philippines)

Serve Robotics • Philippines

On-site
PHP 900,000 - 1,500,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Industrious Ventures • Philippines

On-site
PHP 1,200,000 - 1,800,000
Incident Response Lead
Incident Response Lead

Permhunt • Cebu City

On-site
PHP 900,000 - 1,700,000
Lead Systems & Network Engineer | Onsite
Lead Systems & Network Engineer | Onsite

TerraBarn Inc • Cebu City

On-site
PHP 900,000 - 1,500,000