Robotics Reliability & Incident Response Engineer

Industrious Ventures

Philippines

On-site

PHP 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Serve Robotics is seeking a Reliability Operations Engineer to support the operational reliability of robotic and cloud systems. You will handle Tier 2 escalations, refine runbooks, and perform technical investigations during daytime hours, collaborating with senior engineers, product teams, and SREs to improve system health.

You will lead incident investigations, respond to Tier 1 escalations, update documentation, and run automations to streamline troubleshooting.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent hands-on experience.
  • 5+ years of professional experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support function.
  • Experience participating in Tier 1 or Tier 2 investigations, including log review, basic triage, and structured escalation.
  • Exposure to operational environments supporting distributed or cloud-based systems.
  • Participation in incident response workflows and/or on-call rotations.
  • Proficiency with Linux, including navigating systems, reviewing logs, and performing basic diagnostics.
  • Experience using and contributing to runbooks and operational workflows.
  • Ability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry.

Responsibilities

  • Lead incident investigations during your region’s daytime hours, providing timely updates, escalating appropriately, and supporting senior engineers leading the response.
  • Respond to escalations from Tier 1 support using established runbooks, metrics, logs, and diagnostics to remediate issues or escalation to Tier 3 when needed.
  • Update runbooks and operational documentation based on new issues, discoveries, and feedback, ensuring clarity and consistency across all procedures.
  • Run existing automations and collaborate with senior team members to enhance tooling and scripts that streamline troubleshooting and remediation tasks.
  • Use observability tools such as Grafana/Prometheus, GCP Monitoring, and OpenTelemetry to interpret metrics, logs, and traces, helping identify anomalies and validate system performance.
  • Provide concise, accurate updates during incidents, ensuring information reaches the correct engineering and SRE contacts and supporting structured incident coordination.
  • Participate in discussions around root causes, share operational insights, and contribute to process improvements that enhance system stability and supportability.
  • Participate in a shared weekend on-call rotation to help maintain operational coverage for production systems, responding to incidents and escalations as needed and coordinating with engineering teams when issues arise.
  • Proactively strengthen workflows, adopt best practices, and build the foundation of the Reliability Operations function as it evolves.

Skills

Reliability Ops
DevOps
IT Operations
Incident Response
On-call

Education

Bachelor’s degree in CS/IT/Engineering

Tools

Grafana/Prometheus
Google Cloud Monitoring
OpenTelemetry
Jira

Job description

Serve Robotics is seeking a Reliability Operations Engineer to support the operational reliability of robotic and cloud systems. You will handle Tier 2 escalations, refine runbooks, and perform technical investigations during daytime hours, collaborating with senior engineers, product teams, and SREs to improve system health.

You will lead incident investigations, respond to Tier 1 escalations, update documentation, and run automations to streamline troubleshooting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Robotics Reliability & Incident Response Engineer
Robotics Reliability & Incident Response Engineer

Serve Robotics • Philippines

On-site
PHP 900,000 - 1,500,000
Reliability Operations Engineer (Philippines)
Reliability Operations Engineer (Philippines)

Serve Robotics • Philippines

On-site
PHP 900,000 - 1,500,000
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Industrious Ventures • Philippines

On-site
PHP 1,200,000 - 1,800,000
Service Reliability Engineer
Service Reliability Engineer

Metrobank • Philippines

On-site
PHP 480,000 - 720,000
Remote Incident & Reliability Lead for AI Ops
Remote Incident & Reliability Lead for AI Ops

Ethos • Manila

On-site
PHP 5,765,000 - 7,799,000
IT SRE Team Lead: Automation, SLOs & Incident Response
IT SRE Team Lead: Automation, SLOs & Incident Response

Cerebras Systems • Binangonan

On-site
PHP 1,800,000 - 3,200,000
Robotics Application Engineer — Deploy & Optimize AI Solutions
Robotics Application Engineer — Deploy & Optimize AI Solutions

Sereact • Boston

On-site
PHP 7,697,000 - 12,316,000
Medical, dental, vision insurance for‑
401(k) with company match
20 days paid time off
+5
IoT Telemetry SRE: Reliability, Automation & 24x7 Ops
IoT Telemetry SRE: Reliability, Automation & 24x7 Ops

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc. • Taguig

On-site
PHP 1,000,000 - 1,500,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000