Site Reliability Engineer

Confidential Jobs

Kuala Lumpur

On-site

MYR 150,000 - 230,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Confidential Jobs in Malaysia seeks a Site Reliability Engineer to ensure the reliability, performance, and scalability of cloud infrastructure across Alibaba Cloud (Alicloud) and AWS environments.

You will work closely with engineering, infrastructure, and application teams to build resilient systems, automate operational processes, and improve service availability across the technology landscape.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or related roles.
  • Hands-on experience managing production environments on Alibaba Cloud (Alicloud) and/or AWS.
  • Strong experience with cloud infrastructure operations, monitoring, incident management, and troubleshooting distributed systems.
  • Experience automating operational tasks using scripting, IaC, or configuration management tools.
  • Understanding of reliability, scalability, performance tuning, backup, recovery, and high availability.
  • Experience collaborating with software engineering, infrastructure, and platform teams.
  • Demonstrates ownership, resilience under pressure, and a continuous improvement mindset.

Responsibilities

  • Platform Reliability & Operations: maintain and improve availability, performance, and resilience of cloud infrastructure across Alicloud and AWS; implement automation to reduce effort and improve reliability; establish observability practices; lead incident response and post-incident reviews.

Skills

Cloud engineering
Incident response
Automation
Observability
Collaboration

Tools

Alibaba Cloud
AWS
Monitoring tools

Job description

We are seeking a Site Reliability Engineer to help ensure the reliability, performance, and scalability of our cloud infrastructure and business-critical applications across Alibaba Cloud (Alicloud) and AWS environments. This role will play a key part in enhancing platform stability, automating operational processes, and improving service availability across the technology landscape.

You will work closely with engineering, infrastructure, and application teams to build resilient systems, reduce operational risk, and strengthen observability practices. This is an opportunity to contribute to a modern cloud environment while driving continuous improvement in reliability and operational excellence.

Key Responsibilities
Platform Reliability & Operations
  • Maintain and improve the availability, performance, and resilience of cloud infrastructure across Alibaba Cloud (Alicloud) and AWS through proactive monitoring, incident management, and root cause analysis.
  • Implement automation solutions that reduce operational effort, improve consistency, and strengthen system reliability at scale.
  • Establish and maintain observability practices, including monitoring, alerting, logging, and performance analysis, to identify and resolve issues before they impact users.
  • Lead incident response activities, coordinate cross-functional resolution efforts, and drive continuous improvement through post-incident reviews and preventive actions.
Infrastructure Engineering & Continuous Improvement
  • Design, implement, and optimize cloud infrastructure configurations across Alicloud and AWS that support scalability, security, and operational stability.
  • Collaborate with development and product teams to improve application reliability, deployment processes, and service performance throughout the software lifecycle.
  • Develop and maintain infrastructure-as-code and operational runbooks to ensure consistent, repeatable, and efficient platform management.
  • Identify reliability risks and recommend pragmatic solutions through data-driven analysis, stakeholder collaboration, and strong problem-solving capabilities.
Qualifications & Requirements
  • 5+ years of experience in Site Reliability Engineering, Cloud Infrastructure Engineering, DevOps, or related roles.
  • Hands-on experience managing and supporting production environments on Alibaba Cloud (Alicloud) and/or AWS.
  • Strong experience with cloud infrastructure operations, monitoring tools, incident management, and troubleshooting complex distributed systems.
  • Experience automating operational tasks using scripting, infrastructure-as-code, or configuration management tools.
  • Understanding of system reliability, scalability, performance tuning, backup, recovery, and high-availability principles.
  • Experience working collaboratively with software engineering, infrastructure, and platform teams.
  • Character: Demonstrates ownership, accountability, resilience under pressure, a continuous improvement mindset, strong problem-solving ability, and a collaborative approach to achieving outcomes.
Good-to-Have
  • Professional certifications in AWS and/or Alibaba Cloud.
  • Experience with containerization and orchestration technologies.
  • Experience supporting CI/CD pipelines and release automation.
  • Exposure to multi-cloud or hybrid-cloud environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multi-Cloud Site Reliability Engineer: Observability & Automation
Multi-Cloud Site Reliability Engineer: Observability & Automation

Confidential Jobs • Kuala Lumpur

On-site
MYR 150,000 - 230,000
Senior Alicloud Engineer
Senior Alicloud Engineer

Confidential Jobs • Kuala Lumpur

On-site
MYR 150,000 - 230,000
Senior Devops Engineer
Senior Devops Engineer

MHA Consultancy Services Sdn Bhd • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior Manager, Cloud Architect (Alibaba Cloud)
Senior Manager, Cloud Architect (Alibaba Cloud)

AIA Hong Kong and Macau • Kuala Lumpur

On-site
MYR 150,000 - 200,000
Senior Manager, Cloud Solution Architect (Ali Cloud)
Senior Manager, Cloud Solution Architect (Ali Cloud)

AIA Digital+ Malaysia • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Cloud Operations Engineer (Platform Reliability / NOC)
Cloud Operations Engineer (Platform Reliability / NOC)

Agensi Pekerjaan Genie Hunt Talent • Petaling Jaya

On-site
MYR 60,000 - 80,000
IDC IT Ops Engineer
IDC IT Ops Engineer

ZTE Malaysia • Johor Bahru

On-site
MYR 48,000 - 78,000
IT Operations Analyst
IT Operations Analyst

Accenture Southeast Asia • Cyberjaya

On-site
MYR 120,000 - 180,000
Senior Security Operations lead
Senior Security Operations lead

Randstad Malaysia • Kuala Lumpur

On-site
MYR 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

LAVU TECH SOLUTIONS SDN. BHD. • Petaling Jaya

On-site
MYR 180,000 - 300,000