Senior Reliability Operations Engineer (Malaysia)

Serve Robotics

Penang

On-site

MYR 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Serve Robotics is seeking a Senior Reliability Operations Engineer to lead regional incident response for robotic and cloud systems. You will coordinate investigations, develop runbooks, and drive automation projects in collaboration with product engineering and SRE teams.

You will also own the incident lifecycle, craft clear communications, and improve detection and response through automation and documented processes. Weekend on-call rotation is required.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
  • 5+ years of experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or related field.
  • Experience owning or participating in Tier 2/3 investigations, including triage, log analysis, and structured escalation.
  • Experience supporting distributed systems or cloud-hosted services.
  • Hands-on incident response experience.
  • Strong proficiency with Linux for diagnostics and troubleshooting.
  • Experience writing and maintaining runbooks, automations, and workflows.
  • Ability to read metrics, logs, and traces using Grafana/Prometheus, GCP Monitoring, OpenTelemetry.
  • Familiarity with GCP and basic debugging and permissions.
  • Ability to follow documented procedures and escalate when needed.
  • Understanding of CI/CD and microservice dependencies.
  • Proficiency with Jira or similar incident-tracking platforms.
  • Excellent communication during high-pressure incidents.

Responsibilities

  • Lead regional incident response during daytime hours, coordinating investigations and centralized communication.
  • Respond to escalations from Tier 1 using runbooks, logs, and diagnostics to remediate or escalate.
  • Develop and update runbooks, workflows, and operational documentation across teams.
  • Write and maintain automation scripts to speed remediation and reduce manual work.
  • Use Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry to detect issues and improve monitoring.
  • Act as central communication point during active incidents and route to the right stakeholders.
  • Collaborate with reliability and product teams to share insights and improve processes.
  • Participate in a weekend on-call rotation to maintain production coverage.
  • Establish operational best practices and foundational reliability capabilities.

Skills

Linux
Incident response
Runbook automation
Grafana/Prometheus
Google Cloud Platform
OpenTelemetry
Jira/JSM
PagerDuty
Communication under pressure
Cross-team collaboration

Education

Bachelor’s degree in CS/IT/Engineering

Tools

PagerDuty
Opsgenie
Jira Service Management

Job description

At Serve Robotics, we’re reimagining how things move in cities. Our personable sidewalk robot is our vision for the future. It’s designed to take deliveries away from congested streets, make deliveries available to more people, and benefit local businesses.

The Serve fleet has been delighting merchants, customers, and pedestrians along the way in Los Angeles, Miami, Dallas, Atlanta and Chicago while doing commercial deliveries. We’re looking for talented individuals who will grow robotic deliveries from surprising novelty to efficient ubiquity.

Who We Are

We are tech industry veterans in software, hardware, and design who are pooling our skills to build the future we want to live in. We are solving real-world problems leveraging robotics, machine learning and computer vision, among other disciplines, with a mindful eye towards the end-to-end user experience. Our team is agile, diverse, and driven. We believe that the best way to solve complicated dynamic problems is collaboratively and respectfully.

The Senior Reliability Operations Engineer leads operational reliability by region owning incident response, escalations, and Tier 2 support for robotic and cloud systems. This role drives the creation and improvement of runbooks, automations, and operational processes while coordinating closely with product engineering and SREs. This position serves as the regional incident lead, ensuring issues are resolved efficiently and communicated clearly to all stakeholders.

Responsibilities

Serve as the primary incident lead during your region’s daytime hours, coordinating technical investigations, centralizing communication, and engaging the appropriate engineering and SRE teams when escalation is required.

  • Respond to escalations from Tier 1 support, using runbooks, metrics, logs, and system diagnostics to investigate and remediate issues or determine when escalation to Tier 3 is necessary.

  • Develop and update runbooks, workflows, and operational documentation to ensure consistent and reliable responses to recurring issues, collaborating with product teams to expand coverage over time.

  • Write, maintain, and enhance automation scripts and tools that streamline common remediation steps, improve response times, and reduce manual operational overhead.

  • Use metrics, logs, and tracing tools (Grafana/Prometheus, GCP Monitoring, OpenTelemetry) to proactively identify problems, validate system behavior, and support continuous improvement of detection mechanisms.

  • Act as the central point of communication during active incidents, ensuring timely updates and clear routing to the correct product engineering and SRE stakeholders.

  • Collaborate with reliability and product teams to share insights, recommend improvements, and help refine processes that enhance the stability and operability of our systems.

  • Participate in a shared weekend on-call rotation to help maintain operational coverage for production systems, responding to incidents and escalations as needed and coordinating with engineering teams when issues arise.

  • Help establish operational best practices, refine workflows, and prepare the foundation for a broader reliability operations function.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.

  • 5+ years of professional experience in Reliability Operations, Site Reliability Engineering, DevOps, IT Operations, or a related technical support function.

  • Demonstrated experience owning or participating in Tier 2 or Tier 3 technical investigations, including triage, log analysis, and structured escalation.

  • Experience supporting distributed systems, cloud-hosted services, or production operational environments.

  • Hands-on experience participating in incident response processes.

  • Strong proficiency with Linux, including navigating systems, reviewing logs, and performing diagnostics.

  • Experience writing, executing, and maintaining runbooks, automations, and operational workflows.

  • Ability to interpret metrics, logs, and traces using tools such as Grafana/Prometheus, Google Cloud Monitoring, and OpenTelemetry.

  • Familiarity with modern cloud environments, preferably Google Cloud Platform (GCP), including basic debugging, permissions and service-level triage.

  • Ability to investigate and remediate issues following documented procedures, escalating effectively when needed.

  • Understanding of CI/CD pipelines, deployed application behavior, and operational dependencies across microservices.

  • Proficiency with Jira or similar platforms for ticketing and structured incident tracking.

  • Exceptional communication skills, especially during high-pressure incidents where clear, concise updates are critical.

  • Calm and methodical approach to troubleshooting, prioritization, and decision-making.

  • Strong collaboration skills when coordinating with product engineering, SRE, and global support teams.

  • High level of ownership, reliability and accountability when handling operational responsibilities and incident leadership.

What Makes You Stand Out

  • Experience acting as an incident commander or primary incident response lead for high-severity events.

  • Hands-on experience with robot fleets, IoT devices, or edge systems operations.

  • Experience building lightweight tools, scripts, or internal automations to increase operational efficiency.

  • Familiarity with incident management tools such as PagerDuty, OpsGenie, Jira Service Management, or Grafana IRM.

  • Background creating or improving operational documentation, runbooks, or support processes at scale.

  • Ability to coach and mentor others, and to uplift operational maturity within a region or team.

  • Strong networking fundamentals, including experience diagnosing connectivity issues across distributed systems. Familiarity with Tailscale or similar zero-trust networking tools is a major plus.

Additional Information

  • As part of maintaining continuous operational coverage, this role also participates in a rotating weekend on-call schedule shared across the Reliability Operations team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Seeds Renewables • Kuala Lumpur

On-site
MYR 60,000 - 80,000
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Serve Robotics • Penang

On-site
MYR 485,000 - 607,000
Robotics Reliability & Incident Response Engineer
Robotics Reliability & Incident Response Engineer

Serve Robotics • Penang

On-site
MYR 485,000 - 607,000
Electrical Engineer
Electrical Engineer

Serve Robotics • Penang

On-site
MYR 120,000 - 180,000
Electrical Engineer
Electrical Engineer

Industrious Ventures • Malaysia

On-site
MYR 90,000 - 130,000
Robotics Service Engineer (Mr. Robot Project)
Robotics Service Engineer (Mr. Robot Project)

MR DIY International • Seri Kembangan

On-site
MYR 67,000 - 100,000
SRE Lead
SRE Lead

Chubb • Malaysia

On-site
MYR 300,000 - 420,000
Senior DevOps Engineer
Senior DevOps Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Data Analyst, Supply Chain
Data Analyst, Supply Chain

Embedded Shishya • Malaysia

On-site
MYR 60,000 - 90,000
Regional Site Reliability Engineer (SRE)
Regional Site Reliability Engineer (SRE)

Zuspresso (M) Sdn Bhd • Shah Alam

On-site
MYR 120,000 - 180,000