Senior Site Reliability Engineer

AirAsia rewards

Kuala Lumpur

On-site

MYR 180,000 - 280,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AirAsia rewards seeks a Senior Site Reliability Engineer to join the Platform Engineering team in Kuala Lumpur. You will design, build, and operate highly available cloud platforms, driving automation across the software delivery lifecycle and collaborating with Engineering, Security, DevOps, and Product teams.

The role focuses on GCP-based platform reliability, IaC with Terraform, GitOps with Argo CD, CI/CD with GitLab, and championing observability, incident management, and security practices.

Qualifications

  • Bachelor's degree or equivalent practical experience.
  • 5+ years in Site Reliability/Platform/Cloud Engineering or DevOps.
  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Strong experience with Terraform and Infrastructure as Code.
  • Hands-on experience with GitLab CI/CD.
  • Experience implementing GitOps using Argo CD.
  • Experience managing Cloudflare services (DNS, WAF, CDN, Load Balancing).
  • Linux administration and troubleshooting skills.
  • Experience with container technologies (Docker, Kubernetes).
  • Scripting skills (Bash, Python, or Go).
  • Experience with monitoring, logging, and observability platforms.
  • Experience with incident management, production support, on-call operations.
  • Excellent troubleshooting and root cause analysis, plus communication skills.

Responsibilities

  • Design, build, and operate highly available production platforms on GCP.
  • Develop IaC using Terraform to provision cloud infra.
  • Implement GitOps workflows with Argo CD and GitLab.
  • Build CI/CD pipelines with GitLab for secure, reliable delivery.
  • Develop automation to reduce repetitive tasks via scripting and APIs.
  • Manage Cloudflare services (DNS, WAF, CDN, Load Balancing).
  • Build observability platforms (monitoring, logging, dashboards, SLO/SLI).
  • Drive reliability through monitoring, capacity planning, performance tuning, automation.
  • Participate in on-call rotation, troubleshoot incidents, perform RCAs, implement fixes.
  • Collaborate with development teams to improve reliability, deployment strategies, and readiness.
  • Ensure platform security with best practices, secrets management, least-privilege access.
  • Create/maintain runbooks and documentation.
  • Mentor junior engineers and promote SRE practices.

Skills

SRE practices
Troubleshooting
Root cause analysis
Automation
Stakeholder management
On-call operations

Education

Bachelor's degree in CS/IT/Engineering or equivalent

Tools

GCP
Terraform
GitLab CI/CD
Argo CD
Cloudflare
Kubernetes
Docker

Job description

We are looking for a highly motivated Senior Site Reliability Engineer (SRE) to join our Platform Engineering team. In this role, you will design, build, and operate highly available, secure, and scalable cloud platforms while driving automation across the software delivery lifecycle.

You will partner closely with Engineering, Security, DevOps, and Product teams to improve platform reliability, developer productivity, operational excellence, and cloud governance. This is a hands-on engineering role requiring strong expertise in cloud infrastructure, Infrastructure as Code (IaC), GitOps, observability, and incident management.

Key Responsibilities
  • Design, build, and operate highly available production platforms on Google Cloud Platform (GCP).
  • Develop Infrastructure as Code (IaC) using Terraform to provision and manage cloud infrastructure.
  • Implement and maintain GitOps workflows using Argo CD and GitLab.
  • Build and enhance CI/CD pipelines using GitLab to enable secure, reliable, and automated software delivery.
  • Develop automation solutions to eliminate repetitive operational tasks using scripting and APIs.
  • Manage and optimize Cloudflare services including DNS, WAF, CDN, Load Balancing, Zero Trust, and security controls.
  • Build and maintain observability platforms including monitoring, logging, alerting, tracing, dashboards, and SLO/SLI reporting.
  • Drive platform reliability through proactive monitoring, capacity planning, performance tuning, resilience testing, and automation.
  • Participate in an on-call rotation, troubleshoot production incidents, lead incident response, perform root cause analysis (RCA), and implement permanent corrective actions.
  • Improve operational excellence by reducing toil through automation and self-service capabilities.
  • Collaborate with development teams to improve application reliability, deployment strategies, and operational readiness.
  • Ensure platform security by implementing infrastructure best practices, policy enforcement, secrets management, and least-privilege access.
  • Create and maintain technical documentation, operational runbooks, and standard operating procedures.
  • Mentor junior engineers and promote SRE best practices across engineering teams.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps.
  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Strong experience with Terraform and Infrastructure as Code.
  • Hands-on experience with GitLab CI/CD.
  • Experience implementing GitOps using Argo CD.
  • Experience managing Cloudflare services including DNS, WAF, CDN, and Load Balancing.
  • Strong Linux administration and troubleshooting skills.
  • Experience with container technologies including Docker and Kubernetes.
  • Strong scripting skills using Bash, Python, or Go.
  • Experience with monitoring, logging, and observability platforms.
  • Experience with incident management, production support, and on-call operations.
  • Excellent troubleshooting and root cause analysis skills.Strong communication and stakeholder management skills.
Preferred Qualifications
  • Experience operating Kubernetes platforms such as GKE.
  • Experience with service mesh technologies (Istio, Linkerd, or Envoy).
  • Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and Toil Reduction.
  • Experience implementing platform security and DevSecOps practices.
  • Experience with FinOps and cloud cost optimization.
  • Experience with policy-as-code and infrastructure governance.
  • Google Cloud Professional certifications are an advantage.
  • Knowledge in API's and gateways is added advantages
What Success Looks Like

Within your first 12 months, you will:

  • Improve platform reliability and availability through automation and engineering improvements.
  • Reduce operational toil by automating manual processes.
  • Improve deployment reliability using GitOps and CI/CD best practices.
  • Enhance observability with actionable monitoring and alerting.
  • Strengthen platform security and operational governance.
  • Enable engineering teams to deliver software faster and more reliably.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Career Wise • Kuala Lumpur

On-site
MYR 80,000 - 120,000
SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
SRE Lead
SRE Lead

Chubb Ltd. • Malaysia

On-site
MYR 240,000 - 420,000
Site Reliability Engineer (4024)
Site Reliability Engineer (4024)

Uncover • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AirAsia • Malaysia

On-site
MYR 180,000 - 260,000
Engineering Manager – Platform & SRE
Engineering Manager – Platform & SRE

INSCALE • Kuala Lumpur

On-site
MYR 350,000 - 650,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
Senior DevOps Engineer
Senior DevOps Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Cloud Platform SRE: Reliability, GitOps & Observability
Cloud Platform SRE: Reliability, GitOps & Observability

AirAsia • Malaysia

On-site
MYR 180,000 - 260,000
Regional Site Reliability Engineer (SRE)
Regional Site Reliability Engineer (SRE)

Zuspresso (M) Sdn Bhd • Shah Alam

On-site
MYR 120,000 - 180,000