Reliability Operations Engineer (Philippines)

Serve Robotics

Philippines

On-site

PHP 900,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Serve Robotics is seeking a Reliability Operations Engineer to support robotic and cloud systems. You will lead daytime incident investigations, triage escalations, and refine runbooks while collaborating with senior engineers and SREs.

The role requires 5+ years in reliability operations, strong Linux skills, and hands-on experience with GCP and observability tooling. Weekend on-call rotation is part of the role.

Qualifications

  • Bachelor’s degree in Computer Science, IT, Engineering, or equivalent hands-on experience.
  • 5+ years in Reliability Operations, SRE, DevOps, or related technical support.
  • Experience in Tier 1/2 investigations with log review and escalation.
  • Familiarity with cloud environments and incident response workflows.
  • Proficiency with Linux, logs, and basic diagnostics.
  • Ability to follow runbooks and document remediation steps.

Responsibilities

  • Lead incident investigations during regional daytime hours with timely updates.
  • Respond to Tier 1/2 escalations using runbooks, metrics, and diagnostics.
  • Update runbooks and docs for new issues and discoveries.
  • Run automations and improve tooling to streamline remediation tasks.
  • Use Grafana/Prometheus, GCP Monitoring, and OpenTelemetry to interpret metrics and logs.
  • Provide concise updates during incidents and coordinate with engineers.

Skills

Incident response
Linux
Cloud platforms (GCP)
Runbooks
Observability
Jira
On-call rotations
Collaboration with SREs

Education

Bachelor's degree

Tools

Grafana
Prometheus
Google Cloud Monitoring
OpenTelemetry

Job description

At Serve Robotics, we’re reimagining how things move in cities. Our personable sidewalk robot is our vision for the future. It’s designed to take deliveries away from congested streets, make deliveries available to more people, and benefit local businesses.

The Serve fleet has been delighting merchants, customers, and pedestrians along the way in Los Angeles, Miami, Dallas, Atlanta and Chicago while doing commercial deliveries. We’re looking for talented individuals who will grow robotic deliveries from surprising novelty to efficient ubiquity.

Who We Are

We are tech industry veterans in software, hardware, and design who are pooling our skills to build the future we want to live in. We are solving real-world problems leveraging robotics, machine learning and computer vision, among other disciplines, with a mindful eye towards the end-to-end user experience. Our team is agile, diverse, and driven. We believe that the best way to solve complicated dynamic problems is collaboratively and respectfully.

The Reliability Operations Engineer supports the operational reliability of robotic and cloud systems by handling Tier 2 escalations, following and improving runbooks, and performing technical investigations during your region’s daytime hours. This role works closely with senior team members, product engineering, and SREs to investigate issues, refine operational workflows, and strengthen system health. This position contributes to incident response by providing triage and clear communication, ensuring timely escalation and effective coordination across teams.

Responsibilities

  • Lead incident investigations during your region’s daytime hours, providing timely updates, escalating appropriately, and supporting senior engineers leading the response.

  • Respond to escalations from Tier 1 support using established runbooks, metrics, logs, and diagnostics to remediate issues or escalation to Tier 3 when needed.

  • Update runbooks and operational documentation based on new issues, discoveries, and feedback, ensuring clarity and consistency across all procedures.

  • Run existing automations and collaborate with senior team members to enhance tooling and scripts that streamline troubleshooting and remediation tasks

  • Use observability tools such as Grafana/Prometheus, GCP Monitoring, and OpenTelemetry to interpret metrics, logs, and traces, helping identify anomalies and validate system performance.

  • Provide concise, accurate updates during incidents, ensuring information reaches the correct engineering and SRE contacts and supporting structured incident coordination.

  • Participate in discussions around root causes, share operational insights, and contribute to process improvements that enhance system stability and supportability.

  • Participate in a shared weekend on-call rotation to help maintain operational coverage for production systems, responding to incidents and escalations as needed and coordinating with engineering teams when issues arise.

  • Proactively strengthen workflows, adopt best practices, and build the foundation of the Reliability Operations function as it evolves.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent hands-on experience.

  • 5+ years of professional experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support function.

  • Experience participating in Tier 1 or Tier 2 investigations, including log review, basic triage, and structured escalation.

  • Exposure to operational environments supporting distributed or cloud-based systems.

  • Participation in incident response workflows and/or on-call rotations.

  • Proficiency with Linux, including navigating systems, reviewing logs, and performing basic diagnostics.

  • Experience using and contributing to runbooks and operational workflows.

  • Ability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry.

  • Familiarity with cloud platforms, preferably Google Cloud Platform (GCP).

  • Ability to follow documented remediation steps, with good judgment around when to escale.

  • Understanding of CI/CD pipelines and how application deployments affect runtime behavior.

  • Experience using Jira or similar ticketing systems.

  • Clear and effective communicator, especially when providing updates during time-sensitive operational issues.

  • Calm, organized approach to troubleshooting and prioritization.

  • Collaborative mindset, working effectively with senior operations engineers, product teams, and SREs.

  • Strong sense of ownership and accountability for operational responsibilities.

What Makes You Stand You

  • Prior experience participating in high-severity incident response or supporting operational incidents.

  • Exposure to robot fleets, IoT systems, or other distributed physical device environments.

  • Ability to write or modify lightweight scripts and automations to improve operational workflows.

  • Familiarity with incident management platforms such as PagerDuty, OpsGenie, Jira Service Management, or Grafana IRM.

  • Experience contributing to the creation or improvement of operational runbooks and support documentation.

  • Strong networking fundamentals; familiarity with Tailscale or similar zero-trust networking tools is a plus.

  • Demonstrated ability to learn quickly and contribute to improving operational maturity within a team

Additional Information

  • As part of maintaining continuous operational coverage, this role also participates in a rotating weekend on-call schedule shared across the Reliability Operations team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Industrious Ventures • Philippines

On-site
PHP 1,200,000 - 1,800,000
Robotics Reliability & Incident Response Engineer
Robotics Reliability & Incident Response Engineer

Serve Robotics • Philippines

On-site
PHP 900,000 - 1,500,000
Robotics Reliability & Incident Response Engineer
Robotics Reliability & Incident Response Engineer

Industrious Ventures • Philippines

On-site
PHP 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8, Inc. • Manila

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Private Advertiser • Makati

On-site
PHP 1,200,000 - 1,800,000
Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month
Senior Platform Engineer (DevOps / Site Reliability Engineer) RTO 1x In A Month

AVENSYS CONSULTING INC. • Pasay

Hybrid
PHP 1,000,000 - 1,600,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc. • Taguig

On-site
PHP 1,000,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

8x8, Inc. • Manila

On-site
PHP 781,200 - 1,674,000
Onboarding program
Global-scale production exposure
Blameless post-mortems culture
+1
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000