Robotics Reliability & Incident Response Engineer

Serve Robotics

Penang

On-site

MYR 485,000 - 607,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Serve Robotics is seeking a Reliability Operations Engineer to support the operational reliability of robotic and cloud systems. You will handle Tier 2 escalations during daytime hours, refine runbooks, and perform technical investigations in collaboration with senior engineers, product teams, and SREs.

You will triage incidents, analyze logs and metrics with Grafana/Prometheus and OpenTelemetry, and help improve incident response workflows.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent hands‑on experience.
  • 2–4 years of experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support function.
  • Experience participating in Tier 1 or Tier 2 investigations, including log review, basic triage, and structured escalation.
  • Exposure to operational environments supporting distributed or cloud-based systems.
  • Participation in incident response workflows and/or on‑call rotations.
  • Proficiency with Linux, including navigating systems, reviewing logs, and performing basic diagnostics.
  • Experience using and contributing to runbooks and operational workflows.
  • Ability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry.
  • Familiarity with cloud platforms, preferably Google Cloud Platform (GCP).
  • Ability to follow documented remediation steps, with good judgment around when to escalation.
  • Understanding of CI/CD pipelines and how application deployments affect runtime behavior.
  • Experience using Jira or similar ticketing systems.
  • Clear and effective communicator, especially when providing updates during time-sensitive operational issues.
  • Calm, organized approach to troubleshooting and prioritization.
  • Collaborative mindset, working effectively with senior operations engineers, product teams, and SREs.
  • Strong sense of ownership and accountability for operational responsibilities.

Responsibilities

  • Lead incident investigations during your region’s daytime hours, providing timely updates, escalating appropriately, and supporting senior engineers leading the response.
  • Respond to escalations from Tier 1 support using established runbooks, metrics, logs, and diagnostics to remediate issues or elevate to Tier 3 when needed.
  • Update runbooks and operational documentation based on new issues, discoveries, and feedback, ensuring clarity and consistency across all procedures.
  • Run existing automations and collaborate with senior team members to enhance tooling and scripts that streamline troubleshooting and remediation tasks.
  • Use observability tools such as Grafana/Prometheus, GCP Monitoring, and OpenTelemetry to interpret metrics, logs, and traces, helping identify anomalies and validate system performance.
  • Provide concise, accurate updates during incidents, ensuring information reaches the correct engineering and SRE contacts and supporting structured incident coordination.
  • Participate in discussions around root causes, share operational insights, and contribute to process improvements that enhance system stability and supportability.
  • Participate in a shared weekend on-call rotation to help maintain operational coverage for production systems, responding to incidents and escalations as needed and coordinating with engineering teams when issues arise.
  • Proactively strengthen workflows, adopt best practices, and build the foundation of the Reliability Operations function as it evolves.

Skills

Linux
Incident response
Runbooks
Grafana/Prometheus
OpenTelemetry
GCP
CI/CD
Jira
Communication
On-call
Team collaboration

Education

Bachelor’s degree in CS/IT/Engineering

Tools

Grafana
Prometheus
Google Cloud Monitoring
OpenTelemetry
Jira

Job description

Serve Robotics is seeking a Reliability Operations Engineer to support the operational reliability of robotic and cloud systems. You will handle Tier 2 escalations during daytime hours, refine runbooks, and perform technical investigations in collaboration with senior engineers, product teams, and SREs.

You will triage incidents, analyze logs and metrics with Grafana/Prometheus and OpenTelemetry, and help improve incident response workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Reliability Operations Engineer — Robotics
Senior Reliability Operations Engineer — Robotics

Serve Robotics • Penang

On-site
MYR 180,000 - 240,000
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Seeds Renewables • Kuala Lumpur

On-site
MYR 60,000 - 80,000
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Serve Robotics • Penang

On-site
MYR 485,000 - 607,000
Senior Reliability Operations Engineer (Malaysia)
Senior Reliability Operations Engineer (Malaysia)

Serve Robotics • Penang

On-site
MYR 180,000 - 240,000
Robotics Reliability Ops Engineer
Robotics Reliability Ops Engineer

Seeds Renewables • George Town

On-site
MYR 60,000 - 80,000
Robotics Field Deployment & Integration Engineer
Robotics Field Deployment & Integration Engineer

MR DIY International • Seri Kembangan

On-site
MYR 60,000 - 90,000
SRE Lead
SRE Lead

Chubb Ltd. • Malaysia

On-site
MYR 240,000 - 420,000
Robotics Field Service Engineer
Robotics Field Service Engineer

MR DIY International • Seri Kembangan

On-site
MYR 67,000 - 100,000
SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Career Wise • Kuala Lumpur

On-site
MYR 80,000 - 120,000