Senior Site Reliability Engineer

Carousell Group

Kuala Lumpur

On-site

MYR 180,000 - 300,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mudah.my Sdn. Bhd in Malaysia seeks a seasoned Reliability and Infrastructure Engineer to own production systems at scale. You will manage Kubernetes-based deployments, Linux performance, Docker, and cloud infrastructure using IaC, with Vault for secrets and GitHub Actions for CI/CD.

You will build robust observability with Prometheus and Grafana, implement safe release processes, automate toil, and lead resilience testing across multi-service platforms.

Qualifications

  • 5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.
  • Kubernetes in production deployment, upgrades and troubleshooting.
  • Strong Linux fundamentals and performance tuning plus Docker and container networking.
  • Hands-on Google Cloud Platform experience, with infrastructure as code and declarative provisioning.
  • Secrets and credential management with Vault, rotation and least-privilege access.
  • CI/CD pipelines on GitHub Actions with progressive delivery and rollback.
  • Bash plus Go or Python for tooling.
  • Monitoring with Prometheus and Grafana; cloud instrumentation and dashboards.

Responsibilities

  • Operate against existing SLIs, SLOs and error budgets; balance reliability work with feature velocity.
  • Own incident lifecycle end-to-end: detection, response, mitigation, blameless postmortems.
  • Extend observability stack; add metrics, logs and traces; improve dashboards and alerts.
  • Design and maintain safe release processes: canary, progressive rollout, automated rollback.
  • Automate toil; treat repetitive ops as bugs to engineer away.
  • Manage infrastructure as code provisioning and documentation.
  • Capacity planning, performance tuning and cloud cost efficiency.
  • Security controls, secrets and credential management, access control and monitoring.
  • Conduct production readiness reviews and resilience testing.
  • Operate MCP gateway and AI tooling for production infrastructure and automation.

Skills

SRE
DevOps
Kubernetes
Linux
Docker
GCP
Terraform
Vault
CI/CD
Go
Python
Prometheus
Grafana
GitHub Actions
Google Cloud

Tools

Prometheus
Grafana
GitHub Actions
GCP

Job description

Carousell Group is the leading recommerce group in Greater Southeast Asia on a mission to inspire the world to start selling, and to make secondhand the first choice. Founded in August 2012 in Singapore, the Group has a leading presence in eight markets under the brands Carousell, Cho Tot, Laku6, Mudah.my, OneKyat, Ox Street, and Refash, serving tens of millions of monthly active users. Carousell is backed by leading investors including Telenor Group, Rakuten Ventures, Naver, STIC Investments and Sequoia Capital India.

As a team of passionate individuals working together to solve meaningful problems, there is so much more for you to discover in a career with Carousell. Our culture is made up of hiring, developing, and promoting people who embody our values of solving problems for our users; having a mission-first mindset; being relentlessly resourceful; caring deeply; and staying humble to constantly improve. Together as an organisation, we make magic happen.

About Mudah

Mudah.my Sdn. Bhd is Malaysia’s largest digital platform for selling and finding almost anything - from Cars to Cameras, Properties to Pets, Mobile phones to Motorcycles, Treadmills to Textbooks, Bicycles to Beds, Guitars to Golf sets, Plants to Printers, Watches to Washing machines, Tyres to Tablets, Dresses to Drums, Shoes to Shops, Collectibles to Computers, Jobs and more – Semua Pun Mudah! Mudah’s mission is to democratize commerce by empowering everyone, especially individuals and budding entrepreneurs, with a platform of equal opportunity.

Job Description

Responsibilities

Operate against our existing SLIs, SLOs and error budgets, hold services to them, review and adjust targets as systems evolve, and use budget burn to arbitrate between reliability work and feature velocity.

Own the incident lifecycle end to end detection, response, mitigation, blameless postmortems with tracked follow-through troubleshooting across the whole stack (OS, application, database, cache, network), and mature the on-call rotation around it: alerts tuned for signal, runbooks kept current, MTTD and MTTR trending down.

Extend and improve our observability stack instrument new services, close coverage gaps in metrics, logs and traces, and raise dashboard and alert quality so teams can diagnose their own services.

Design and maintain safe release processes: canary, progressive rollout, automated rollback across dev, staging and production.

Eliminate toil through automation; treat repetitive manual operations as bugs to be engineered away.

Own infrastructure as code provisioning, configuration, policy, and the documentation around it.

Capacity planning, performance tuning and cloud cost efficiency forecast growth, model headroom, upgrade before saturation.

Implement and maintain infrastructure security controls secrets and credential management in Vault, access control, monitoring and response.

Conduct production readiness reviews and resilience testing for new and existing services.

Operate our MCP gateway and internal AI tooling as production infrastructure availability, access control, rate limiting, cost and usage visibility and apply AI-assisted automation to operational work such as incident triage, log summarisation and runbook generation.

Qualifications

Reliability and infrastructure

5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.

Kubernetes in production deployment, upgrades, troubleshooting and ongoing maintenance.

Strong Linux fundamentals and performance tuning (RHEL / CentOS / Debian / Ubuntu), plus Docker and container networking.

Hands-on Google Cloud Platform experience, with infrastructure as code and declarative provisioning (Terraform or equivalent).

Secrets and credential management with HashiCorp Vault (or equivalent) policies, rotation and least-privilege access.

CI/CD delivery pipelines on GitHub Actions (or equivalent), including progressive delivery and automated rollback.

Bash plus working proficiency in a general-purpose language (Go, Python or similar) for building real tooling.

Monitoring and observability with Prometheus, Grafana and Google Cloud Operations instrumentation, dashboards, alerting and tracing.

Able to work independently on large, complex projects with minimal guidance.

AI-assisted operations

Practical experience applying AI and LLM-based tooling to engineering or operational workflows, with a clear view of where it helps and where it doesn't.

Working familiarity with MCP (Model Context Protocol) and agent harnesses tool-calling loops, context management, guardrails and failure handling or the appetite and fundamentals to pick them up quickly.

Additional Information

Why Join Us?

At Mudah, we are evolving towards an AI-first engineering organisation. This role is an opportunity to go beyond traditional technical leadership and help shape how engineering teams build software with AI and agentic workflows.

You will have the opportunity to influence both the technology we build and the way we build it.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mudah.my • Kuala Lumpur

On-site
MYR 180,000 - 320,000
Senior Backend Developer
Senior Backend Developer

Carousell Group • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior Backend Developer
Senior Backend Developer

Carousell • Kuala Lumpur

On-site
MYR 180,000 - 320,000
Senior SRE: AI-Driven Infra, Cloud & Reliability
Senior SRE: AI-Driven Infra, Cloud & Reliability

Mudah.my • Kuala Lumpur

On-site
MYR 180,000 - 320,000
AI Site Reliability Engineer
AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

SGS (Malaysia) Sdn Bhd • Kuching

On-site
MYR 120,000 - 180,000
Senior Site Reliability Engineer - Scale-Up Platform
Senior Site Reliability Engineer - Scale-Up Platform

Aisling Group • Kuala Lumpur

On-site
Site Reliability Engineer
Site Reliability Engineer

Aisling Group • Kuala Lumpur

On-site
Product Designer
Product Designer

Mudah.my • Kuala Lumpur

On-site
MYR 67,000 - 112,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Engg • Kuala Lumpur

On-site
MYR 180,000 - 300,000