Senior Site Reliability Engineer

Mudah.my

Kuala Lumpur

On-site

MYR 180,000 - 320,000

Full time

36 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mudah.my is seeking a seasoned Site Reliability Engineer to join our cloud and infrastructure team in Kuala Lumpur. You will manage production systems at scale, champion reliability, and automate complex operations across the stack.

The role emphasizes observability, safe release processes, and AI-assisted automation to improve incident triage and runbook generation. Excellent Linux and cloud skills are essential.

Qualifications

  • 5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.
  • Kubernetes in production deployment, upgrades, troubleshooting and ongoing maintenance.
  • Strong Linux fundamentals and performance tuning, plus Docker and container networking.
  • Google Cloud Platform experience with IaC and declarative provisioning (Terraform or equivalent).
  • Secrets management with Vault; CI/CD pipelines on GitHub Actions; Bash and Go/Python for tooling.

Responsibilities

  • Operate against SLIs/SLOs and error budgets; arbitrate reliability work vs. feature velocity.
  • Own the incident lifecycle end-to-end, with blameless postmortems and follow-up.
  • Extend observability: add metrics, logs, traces; improve dashboards and alerts.
  • Design and maintain safe release processes: canary, progressive rollout, automated rollback.
  • Automate repetitive operations to reduce toil and manual work.
  • Own infrastructure as code provisioning, configuration, policy, and documentation.
  • Capacity planning, performance tuning and cloud cost forecasting; model headroom and upgrades.
  • Implement security controls, secrets management, access control and monitoring.
  • Conduct production readiness reviews and resilience testing for new/existing services.
  • Operate MCP gateway and internal AI tooling for production infrastructure, cost visibility and runbook automation.

Skills

Kubernetes
Linux
Docker
Go
Python
CI/CD
Vault
Prometheus
Grafana

Tools

Terraform
GitHub Actions

Job description

Company Description

Carousell Group is the leading recommerce group in Greater Southeast Asia on a mission to inspire the world to start selling, and to make secondhand the first choice. Founded in August 2012 in Singapore, the Group has a leading presence in eight markets under the brands Carousell, Cho Tot, Laku6, Mudah.my, OneKyat, Ox Street, and Refash, serving tens of millions of monthly active users. Carousell is backed by leading investors including Telenor Group, Rakuten Ventures, Naver, STIC Investments and Sequoia Capital India.

Company Description

Carousell Group is the leading recommerce group in Greater Southeast Asia on a mission to inspire the world to start selling, and to make secondhand the first choice. Founded in August 2012 in Singapore, the Group has a leading presence in eight markets under the brands Carousell, Cho Tot, Laku6, Mudah.my, OneKyat, Ox Street, and Refash, serving tens of millions of monthly active users. Carousell is backed by leading investors including Telenor Group, Rakuten Ventures, Naver, STIC Investments and Sequoia Capital India. As a team of passionate individuals working together to solve meaningful problems, there is so much more for you to discover in a career with Carousell. Our culture is made up of hiring, developing, and promoting people who embody our values of solving problems for our users; having a mission‑first mindset; being relentlessly resourceful; caring deeply; and staying humble to constantly improve. Together as an organisation, we make magic happen.

About Mudah

Mudah.my Sdn. Bhd is Malaysia’s largest digital platform for selling and finding almost anything - from Cars to Cameras, Properties to Pets, Mobile phones to Motorcycles, Treadmills to Textbooks, Bicycles to Beds, Guitars to Golf sets, Plants to Printers, Watches to Washing machines, Tyres to Tablets, Dresses to Drums, Shoes to Shops, Collectibles to Computers, Jobs and more – Semua Pun Mudah! Mudah’s mission is to democratize commerce by empowering everyone, especially individuals and budding entrepreneurs, with a platform of equal opportunity.

Job Description
  • Operate against our existing SLIs, SLOs and error budgets, hold services to them, review and adjust targets as systems evolve, and use budget burn to arbitrate between reliability work and feature velocity.
  • Own the incident lifecycle end to end detection, response, mitigation, blameless postmortems with tracked follow‑through troubleshooting across the whole stack (OS, application, database, cache, network), and mature the on‑call rotation around it: alerts tuned for signal, runbooks kept current, MTTD and MTTR trending down.
  • Extend and improve our observability stack instrument new services, close coverage gaps in metrics, logs and traces, and raise dashboard and alert quality so teams can diagnose their own services.
  • Design and maintain safe release processes: canary, progressive rollout, automated rollback across dev, staging and production.
  • Eliminate toil through automation; treat repetitive manual operations as bugs to be engineered away.
  • Own infrastructure as code provisioning, configuration, policy, and the documentation around it.
  • Capacity planning, performance tuning and cloud cost efficiency forecast growth, model headroom, upgrade before saturation.
  • Implement and maintain infrastructure security controls secrets and credential management in Vault, access control, monitoring and response.
  • Conduct production readiness reviews and resilience testing for new and existing services.
  • Operate our MCP gateway and internal AI tooling as production infrastructure availability, access control, rate limiting, cost and usage visibility and apply AI‑assisted automation to operational work such as incident triage, log summarisation and runbook generation.
Qualifications
Reliability and infrastructure
  • 5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.
  • Kubernetes in production deployment, upgrades, troubleshooting and ongoing maintenance.
  • Strong Linux fundamentals and performance tuning (RHEL / CentOS / Debian / Ubuntu), plus Docker and container networking.
  • Hands‑on Google Cloud Platform experience, with infrastructure as code and declarative provisioning (Terraform or equivalent).
  • Secrets and credential management with HashiCorp Vault (or equivalent) policies, rotation and least‑privilege access.
  • CI/CD delivery pipelines on GitHub Actions (or equivalent), including progressive delivery and automated rollback.
  • Bash plus working proficiency in a general‑purpose language (Go, Python or similar) for building real tooling.
  • Monitoring and observability with Prometheus, Grafana and Google Cloud Operations instrumentation, dashboards, alerting and tracing.
  • Able to work independently on large, complex projects with minimal guidance.
AI‑assisted operations
  • Practical experience applying AI and LLM‑based tooling to engineering or operational workflows, with a clear view of where it helps and where it doesn't.
  • Working familiarity with MCP (Model Context Protocol) and agent harnesses tool‑calling loops, context management, guardrails and failure handling or the appetite and fundamentals to pick them up quickly.
Why Join Us

At Mudah, we are evolving towards an AI‑first engineering organisation. This role is an opportunity to go beyond traditional technical leadership and help shape how engineering teams build software with AI and agentic workflows. You will have the opportunity to influence both the technology we build and the way we build it. you are adhering to our PDPA policies. In case you are interested to know more, read about our Candidates Personal Data Privacy Statement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Carousell Group • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior Backend Developer
Senior Backend Developer

Carousell Group • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior Backend Developer
Senior Backend Developer

Carousell • Kuala Lumpur

On-site
MYR 180,000 - 320,000
Product Designer
Product Designer

Mudah.my • Kuala Lumpur

On-site
MYR 67,000 - 112,000
Product Designer
Product Designer

Carousell Group • Kuala Lumpur

On-site
MYR 60,000 - 90,000
Product Designer
Product Designer

Mudah • Kuala Lumpur

On-site
MYR 60,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Engg • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior SRE: AI-Driven Infra, Cloud & Reliability
Senior SRE: AI-Driven Infra, Cloud & Reliability

Mudah.my • Kuala Lumpur

On-site
MYR 180,000 - 320,000
AI Site Reliability Engineer
AI Site Reliability Engineer

YTL AI Labs • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Commercial Strategy & Operations Lead
Commercial Strategy & Operations Lead

Carousell • Kuala Lumpur

On-site
MYR 180,000 - 300,000