Senior SRE: AI-Driven Reliability & Cloud Automation

Carousell Group

Kuala Lumpur

On-site

MYR 180,000 - 300,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mudah.my Sdn. Bhd in Malaysia seeks a seasoned Reliability and Infrastructure Engineer to own production systems at scale. You will manage Kubernetes-based deployments, Linux performance, Docker, and cloud infrastructure using IaC, with Vault for secrets and GitHub Actions for CI/CD.

You will build robust observability with Prometheus and Grafana, implement safe release processes, automate toil, and lead resilience testing across multi-service platforms.

Qualifications

  • 5+ years operating production systems at scale in SRE, DevOps or infrastructure engineering.
  • Kubernetes in production deployment, upgrades and troubleshooting.
  • Strong Linux fundamentals and performance tuning plus Docker and container networking.
  • Hands-on Google Cloud Platform experience, with infrastructure as code and declarative provisioning.
  • Secrets and credential management with Vault, rotation and least-privilege access.
  • CI/CD pipelines on GitHub Actions with progressive delivery and rollback.
  • Bash plus Go or Python for tooling.
  • Monitoring with Prometheus and Grafana; cloud instrumentation and dashboards.

Responsibilities

  • Operate against existing SLIs, SLOs and error budgets; balance reliability work with feature velocity.
  • Own incident lifecycle end-to-end: detection, response, mitigation, blameless postmortems.
  • Extend observability stack; add metrics, logs and traces; improve dashboards and alerts.
  • Design and maintain safe release processes: canary, progressive rollout, automated rollback.
  • Automate toil; treat repetitive ops as bugs to engineer away.
  • Manage infrastructure as code provisioning and documentation.
  • Capacity planning, performance tuning and cloud cost efficiency.
  • Security controls, secrets and credential management, access control and monitoring.
  • Conduct production readiness reviews and resilience testing.
  • Operate MCP gateway and AI tooling for production infrastructure and automation.

Skills

SRE
DevOps
Kubernetes
Linux
Docker
GCP
Terraform
Vault
CI/CD
Go
Python
Prometheus
Grafana
GitHub Actions
Google Cloud

Tools

Prometheus
Grafana
GitHub Actions
GCP

Job description

Mudah.my Sdn. Bhd in Malaysia seeks a seasoned Reliability and Infrastructure Engineer to own production systems at scale. You will manage Kubernetes-based deployments, Linux performance, Docker, and cloud infrastructure using IaC, with Vault for secrets and GitHub Actions for CI/CD.

You will build robust observability with Prometheus and Grafana, implement safe release processes, automate toil, and lead resilience testing across multi-service platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI-Driven Infra, Cloud & Reliability
Senior SRE: AI-Driven Infra, Cloud & Reliability

Mudah.my • Kuala Lumpur

On-site
MYR 180,000 - 320,000
Senior Site Reliability Engineer: Scale, Resilience & Automation
Senior Site Reliability Engineer: Scale, Resilience & Automation

LAVU TECH SOLUTIONS SDN. BHD. • Petaling Jaya

On-site
MYR 180,000 - 300,000
Senior SRE: Automation, Reliability & Scale
Senior SRE: Automation, Reliability & Scale

Swift Software • Kuala Lumpur

On-site
MYR 120,000 - 240,000
Senior SRE: Cloud, Kubernetes & Automation Leader
Senior SRE: Cloud, Kubernetes & Automation Leader

Apply • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Senior DevOps Engineer — Production Reliability & Automation
Senior DevOps Engineer — Production Reliability & Automation

Cognizant • Kuala Lumpur

On-site
MYR 120,000 - 180,000
SRE Lead: Architect Reliability, Observability & Automation
SRE Lead: Architect Reliability, Observability & Automation

Chubb • Malaysia

On-site
MYR 300,000 - 420,000
Cloud-Native Site Reliability Engineer
Cloud-Native Site Reliability Engineer

SGS (Malaysia) Sdn Bhd • Kuching

On-site
MYR 120,000 - 180,000
SRE Engineer (DevOps) – Global Cloud Reliability
SRE Engineer (DevOps) – Global Cloud Reliability

Ant International • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Site Reliability Engineer: Build Reliable, Scalable Systems
Site Reliability Engineer: Build Reliable, Scalable Systems

Setel Ventures • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Leisure area with video games
Casual dress (jeans)
Pantry with coffee, tea and snacks
+2
Senior DevOps Engineer - Kubernetes, CI/CD & Observability
Senior DevOps Engineer - Kubernetes, CI/CD & Observability

Lenovo • Kuala Lumpur

On-site
MYR 120,000 - 180,000