Senior Site Reliability Engineer

AirAsia

Malaysia

On-site

MYR 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AirAsia Berhad seeks a Senior Site Reliability Engineer to join the Platform Engineering team. You will design, build, and operate highly available cloud platforms, driving automation across the software delivery lifecycle.

You will collaborate with Engineering, Security, DevOps, and Product teams to improve platform reliability, developer productivity, operational excellence, and cloud governance, with hands-on work on GCP, IaC, GitOps, observability, and incident management.

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or equivalent practical experience.
  • 5+ years in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps.
  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Strong experience with Terraform and Infrastructure as Code.
  • Hands-on experience with GitLab CI/CD.
  • Experience implementing GitOps using Argo CD.
  • Experience managing Cloudflare services including DNS, WAF, CDN, and Load Balancing.
  • Strong Linux administration and troubleshooting skills.
  • Experience with container technologies including Docker and Kubernetes.
  • Strong scripting skills using Bash, Python, or Go.
  • Experience with monitoring, logging, and observability platforms.
  • Experience with incident management, production support, and on-call operations.
  • Excellent troubleshooting and root cause analysis skills.
  • Strong communication and stakeholder management skills.

Responsibilities

  • Design, build, and operate highly available production platforms on Google Cloud Platform (GCP).
  • Develop Infrastructure as Code (IaC) using Terraform to provision and manage cloud infrastructure.
  • Implement and maintain GitOps workflows using Argo CD and GitLab.
  • Build and enhance CI/CD pipelines using GitLab to enable secure, reliable, and automated software delivery.
  • Develop automation solutions to eliminate repetitive operational tasks using scripting and APIs.
  • Manage and optimize Cloudflare services including DNS, WAF, CDN, Load Balancing, Zero Trust, and security controls.
  • Build and maintain observability platforms including monitoring, logging, alerting, tracing, dashboards, and SLO/SLI reporting.
  • Drive platform reliability through proactive monitoring, capacity planning, performance tuning, resilience testing, and automation.
  • Participate in an on-call rotation, troubleshoot production incidents, lead incident response, perform root cause analysis (RCA), and implement permanent corrective actions.
  • Improve operational excellence by reducing toil through automation and self-service capabilities.
  • Collaborate with development teams to improve application reliability, deployment strategies, and operational readiness.
  • Ensure platform security by implementing infrastructure best practices, policy enforcement, secrets management, and least-privilege access.
  • Create and maintain technical documentation, operational runbooks, and standard operating procedures.
  • Mentor junior engineers and promote SRE best practices across engineering teams.

Skills

GCP
Terraform
GitLab CI/CD
Argo CD
Cloudflare
Linux administration
Docker
Kubernetes
Bash/Python/Go scripting
Incident management

Education

Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience

Tools

Terraform
GitLab
Argo CD
Docker
Kubernetes

Job description

Job Description We are looking for a highly motivated Senior Site Reliability Engineer (SRE) to join our Platform Engineering team. In this role, you will design, build, and operate highly available, secure, and scalable cloud platforms while driving automation across the software delivery lifecycle. You will partner closely with Engineering, Security, DevOps, and Product teams to improve platform reliability, developer productivity, operational excellence, and cloud governance. This is a hands-on engineering role requiring strong expertise in cloud infrastructure, Infrastructure as Code (IaC), GitOps, observability, and incident management.

Key Responsibilities
  • Design, build, and operate highly available production platforms on Google Cloud Platform (GCP).
  • Develop Infrastructure as Code (IaC) using Terraform to provision and manage cloud infrastructure.
  • Implement and maintain GitOps workflows using Argo CD and GitLab.
  • Build and enhance CI/CD pipelines using GitLab to enable secure, reliable, and automated software delivery.
  • Develop automation solutions to eliminate repetitive operational tasks using scripting and APIs.
  • Manage and optimize Cloudflare services including DNS, WAF, CDN, Load Balancing, Zero Trust, and security controls.
  • Build and maintain observability platforms including monitoring, logging, alerting, tracing, dashboards, and SLO/SLI reporting.
  • Drive platform reliability through proactive monitoring, capacity planning, performance tuning, resilience testing, and automation.
  • Participate in an on-call rotation, troubleshoot production incidents, lead incident response, perform root cause analysis (RCA), and implement permanent corrective actions.
  • Improve operational excellence by reducing toil through automation and self-service capabilities.
  • Collaborate with development teams to improve application reliability, deployment strategies, and operational readiness.
  • Ensure platform security by implementing infrastructure best practices, policy enforcement, secrets management, and least-privilege access.
  • Create and maintain technical documentation, operational runbooks, and standard operating procedures.
  • Mentor junior engineers and promote SRE best practices across engineering teams.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps.
  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Strong experience with Terraform and Infrastructure as Code.
  • Hands-on experience with GitLab CI/CD.
  • Experience implementing GitOps using Argo CD.
  • Experience managing Cloudflare services including DNS, WAF, CDN, and Load Balancing.
  • Strong Linux administration and troubleshooting skills.
  • Experience with container technologies including Docker and Kubernetes.
  • Strong scripting skills using Bash, Python, or Go.
  • Experience with monitoring, logging, and observability platforms.
  • Experience with incident management, production support, and on-call operations.
  • Excellent troubleshooting and root cause analysis skills.
  • Strong communication and stakeholder management skills.
Preferred Qualifications
  • Experience operating Kubernetes platforms such as GKE.
  • Experience with service mesh technologies (Istio, Linkerd, or Envoy).
  • Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and Toil Reduction.
  • Experience implementing platform security and DevSecOps practices.
  • Experience with FinOps and cloud cost optimization.
  • Experience with policy-as-code and infrastructure governance.
  • Google Cloud Professional certifications are an advantage.
  • Knowledge in API's and gateways is added advantages.
What Success Looks Like

Within your first 12 months, you will:

  • Improve platform reliability and availability through automation and engineering improvements.
  • Reduce operational toil by automating manual processes.
  • Improve deployment reliability using GitOps and CI/CD best practices.
  • Enhance observability with actionable monitoring and alerting.
  • Strengthen platform security and operational governance.
  • Enable engineering teams to deliver software faster and more reliably.
AirAsia Berhad

Asia's leading airline was established with the dream of making flying possible for everyone. Since 2001, AirAsia has swiftly broken travel norms around the globe and has risen to become the world's best. Driven by the Dare to Dream spirit, we pride ourselves in being the region's largest low-cost carrier, serving 24 countries and over 130 destinations. We're not confined by walls, except when we need to answer the call of nature, so all departments mingle every day. As we embrace new technology to become a digital airline, services like BIG Duty Free, BIG Pay, BIG Loyalty, Touristly, ROKKI and Xcite Inflight Entertainment will be an exciting evolution, placing us ahead of the game.

Are you in?

AirAsia is set to take low-cost flying to an all new high with our belief, "Now Everyone Can Fly".

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer II
Software Engineer II

AirAsia • Malaysia

On-site
MYR 120,000 - 180,000
Senior Site Reliability Engineer — Cloud Platform Lead
Senior Site Reliability Engineer — Cloud Platform Lead

AirAsia rewards • Kuala Lumpur

On-site
MYR 180,000 - 280,000
System Support Engineer
System Support Engineer

AirAsia • Malaysia

On-site
MYR 78,000 - 123,000
Cloud Platform SRE: Reliability, GitOps & Observability
Cloud Platform SRE: Reliability, GitOps & Observability

AirAsia • Malaysia

On-site
MYR 180,000 - 260,000
Data Engineer II
Data Engineer II

AirAsia • Kuala Lumpur

On-site
MYR 50,000 - 75,000
Quality Assurance Inspector
Quality Assurance Inspector

AirAsia • Kuala Lumpur

On-site
MYR 67,000 - 100,000
Senior Data Engineer
Senior Data Engineer

AirAsia • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Manager, Company Secretarial
Manager, Company Secretarial

AirAsia • Kuala Lumpur

On-site
MYR 480,000 - 720,000
Medical insurance
Maternity expenses
Flexible work arrangement
+8
Station Manager, Kuala Lumpur
Station Manager, Kuala Lumpur

AirAsia • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Free flights
Staff travel benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AirAsia rewards • Kuala Lumpur

On-site
MYR 180,000 - 280,000