Site Reliability Engineer

GoTo Meeting

Manila

On-site

PHP 558,000 - 781,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Dedicated onboarding and shadow period
Global-scale production exposure
Team culture valuing operational discipline

Job summary

8x8 is seeking a Site Reliability Engineer to ensure platform reliability across its global Unified Communications infrastructure. The role involves incident response, complex issue resolution, and cross-team collaboration.

Ideal candidates will have at least 3 years of relevant experience, strong Linux administration skills, and familiarity with cloud environments. The position is Hybrid, requiring on-site attendance on Tuesdays and Wednesdays.

Competitive perks and a culture focused on operational discipline and automation are offered.

Qualifications

  • 3+ years in a site reliability, platform operations, or infrastructure engineering role.
  • Solid Linux systems administration experience, including network diagnostics.
  • Hands-on experience with major cloud providers like OCI, AWS, or GCP.
  • Calm under pressure with effective communication during incidents.
  • Experience with scripting to automate tasks and parse logs.
  • Structured thinking & strong incident response discipline.
  • Familiarity with SRE concepts such as SLIs and SLOs.

Responsibilities

  • Own platform reliability and incident response for UC infrastructure.
  • Triage and resolve complex issues including service restarts and infrastructure failures.
  • Collaborate with cross-functional teams and communicate production events.
  • Build automations to eliminate manual tasks and address security issues.
  • Participate in sprints to improve tooling and deliver essential infrastructure initiatives.

Skills

Site reliability engineering
Linux systems administration
Cloud provider experience
On-call experience
Scripting in Python or Bash
Incident response discipline
Familiarity with SRE concepts
AI-forward mindset

Tools

Prometheus
Grafana
PagerDuty
Oracle Cloud Infrastructure
Ansible

Job description

8x8 connects our customers and teams globally, empowering CX leaders with performance and insights to make smarter decisions, delight customers, and drive lasting business impact.

About 8x8 UC Operations

The UC Operations team manages the production infrastructure behind 8x8's Unified Communications platform — voice, fax, messaging, and collaboration services used by enterprise customers globally. The team oversees dozens of applications running across more than two thousand service instances worldwide, spanning VoIP infrastructure, messaging brokers, storage systems, and cloud workloads across Oracle Cloud Infrastructure and physical datacenters.

UC Ops sits at the operational center of 8x8 — taking escalations from the NOC, coordinating with Engineering, and working alongside Support, Sales, and Professional Services. The work is complex, the systems are live, and the stakes are real. We are actively moving from reactive operations to a proactive, automation-first SRE model — and we are looking for engineers who want to help build that, not just maintain the status quo.

What You'll Do
  • Production Operations & Incident Response
  • Own platform reliability across global UC infrastructure, driving incident response rather than just resolving in isolation.
  • Triage and resolve complex issues — service restarts, hung processes, infrastructure failures — and act as an escalation for the NOC when frontline teams hit their limit.
  • Execute the unglamorous but essential work: scheduled maintenance, certificate renewals, log rotation — the stuff that prevents failure before it happens.
  • Lead blameless post‑mortems that produce real follow‑through, not action items that disappear into a backlog.
Cross‑Team Collaboration
  • Work directly with Support, Sales, Sales Engineering, NOC, Professional Services, and Engineering teams across 8x8 — this team sits at the operational center of the company.
  • Translate production events into clear, business‑readable communication under pressure; stakeholders across the org depend on your judgment during incidents.
  • Feed operational insight back into engineering — turning recurring failures and patterns into actionable bug reports and platform improvements.
  • Reliability Engineering & Automation: identify recurring manual work and build automation to eliminate it — we treat toil as a bug, not a requirement.
  • Participate in 2‑week sprint cycles to deliver automation, tooling improvements, runbook development, and infrastructure initiatives from a structured backlog.
  • Address security issues as they arise — CVEs, misconfigurations, access control gaps — treated as first‑class work alongside incident response.
  • Define and track SLI, SLO, and SLA to drive honest, data‑driven conversations about where reliability investment is needed.
  • Build and maintain dashboards (Grafana, OCI Log Analytics) that give the team genuine signal; tune alerting to eliminate noise.
  • Leverage AI‑powered tooling to accelerate diagnostics and reduce cognitive load at scale.
On‑Call & Coverage
  • Shared on‑call rotation, approximately 1 week per month — same expectation for every engineer on the team.
  • Escalation is always an option and is encouraged; you are expected to drive the response and know when to pull others in, not to hero it alone.
  • Tooling: PagerDuty for alerting, Jira for tracking, OCI Log Analytics and Grafana for diagnostics.
What We're Looking For

Required

  • 3+ years in a site reliability, platform operations, or infrastructure engineering role — you have run production systems and know what that actually means.
  • Solid Linux systems administration: multi‑service distributed systems, log reading, systemctl, network diagnostics, no GUI required.
  • Hands‑on experience with at least one major cloud provider (OCI, AWS, GCP, or Azure) — compute, storage, IAM, networking fundamentals.
  • On‑call experience: calm under pressure, fast triage, clear communication during an incident.
  • Scripting in Python or Bash — enough to automate a task, parse logs, or hit an API independently.
  • Strong incident response discipline: structured thinking, stakeholder communication, post‑mortems that actually say something.
  • Familiarity with SRE concepts: SLIs, SLOs, error budgets, toil measurement.
  • AI‑forward mindset — you use AI tools as a core part of how you work, not as a novelty.

Preferred

  • Experience with Oracle Cloud Infrastructure (OCI) — compute, networking, Log Analytics, Object Storage.
  • Familiarity with VoIP and SIP infrastructure — registration, trunking, call signaling; this is a UC platform and that knowledge matters.
  • Knowledge of observability tooling: Prometheus, Grafana, PagerDuty, OCI Log Analytics.
  • Experience with Ansible for configuration management and deployment automation.
  • Exposure to infrastructure migrations at scale in multi‑tenant SaaS environments.
What We Offer
  • Dedicated onboarding and shadow period before solo on‑call responsibilities.
  • Direct exposure to global‑scale production infrastructure serving global enterprise customers.
  • A team culture that values operational discipline, blameless post‑mortems, and investing in automation over accepting toil.

Work Arrangement: Hybrid (On‑site Tuesdays and Wednesdays)

Office Location: BGC, Taguig

Shift: US Business Hours

8x8 is proud to provide equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability or genetics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

8x8, Inc. • Manila

On-site
PHP 781,200 - 1,674,000
Onboarding program
Global-scale production exposure
Blameless post-mortems culture
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila

On-site
PHP 900,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8 • Manila

On-site
PHP 2,400,000 - 3,600,000
Hybrid SRE - Automation & Incident Response
Hybrid SRE - Automation & Incident Response

8x8, Inc. • Manila

Hybrid
PHP 781,200 - 1,674,000
Onboarding program
Global-scale production exposure
Blameless post-mortems culture
+1
Senior Site Reliability Engineer - Automation Leader
Senior Site Reliability Engineer - Automation Leader

8x8 • Manila

On-site
PHP 2,400,000 - 3,600,000
Hybrid SRE: Global UC Platform Reliability & Automation
Hybrid SRE: Global UC Platform Reliability & Automation

GoTo Meeting • Manila

Hybrid
Dedicated onboarding and shadow period
Global-scale production exposure
Team culture valuing operational discipline
Technical Support Engineer 3
Technical Support Engineer 3

8x8 International Philippine Branch Office • Manila

On-site
PHP 500,000 - 700,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Reliability Operations Engineer (Philippines)
Reliability Operations Engineer (Philippines)

Serve Robotics • Philippines

On-site
PHP 900,000 - 1,500,000
Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month
Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month

AVENSYS CONSULTING INC. • Pasay

Hybrid
PHP 1,000,000 - 1,600,000