System Reliability Engineer, Consultant

AIA Malaysia

Kuala Lumpur

On-site

MYR 70,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

High-impact team environment
Opportunities for innovation
Influence engineering culture

Job summary

A leading insurance company in Kuala Lumpur seeks a passionate System / Site Reliability Engineer to ensure the reliability, scalability, and performance of its enterprise systems. The ideal candidate will have 3-5 years of experience and a strong background in automation, cloud technologies, and monitoring tools. This is an excellent opportunity to be part of a high-impact team and influence engineering culture through best practices. Enjoy substantial growth and development opportunities as you help improve system resilience.

Qualifications

  • 3-5 years of experience in SRE, DevOps, or Software Engineering roles.
  • Experience supporting front-end applications in production environments.

Responsibilities

  • Ensure system reliability and availability.
  • Monitor application performance and report issues.
  • Automate operational tasks and build internal tools.
  • Collaborate closely with development and infrastructure teams.

Skills

Front-end performance monitoring
Automation
Scripting in Python
Cloud technologies (AWS or Azure)
Docker and Kubernetes

Education

Bachelor's degree in Computer Science, Software Engineering, IT, or related fields

Tools

Dynatrace
Grafana
Prometheus
Elastic Stack
Splunk

Job description

At AIA we’ve started an exciting movement to create a healthier, more sustainable future for everyone. As pioneering innovators for over 100 years, we’re now transforming our organisation to be faster, simpler and more connected. Because we want to be even better equipped to develop digital solutions and experiences that help more people live Healthier, Longer, Better Lives. To get there, we need people with tech/digital/analytics expertise and passion to help develop positive, sustainable change through digitally enhanced experiences that will impact the lives of millions of people and create a healthier future for everyone. If you believe in developing a better tomorrow, read on.

About The Role

We are looking for a System / Site Reliability Engineer (SRE) to help ensure the reliability, scalability, and performance of our enterprise systems and services. In this role, you will apply software engineering principles to operations, partner closely with development and infrastructure teams, and build automation that strengthens system stability and efficiency. You will play a pivotal role in bridging the gap between software development and IT operations, driving a culture of resilience, observability, automation, and proactive problem‑solving.

Key Responsibilities
  • Ensure System Reliability & Availability
  • Monitor and report on application performance, and highlight any deviations or issues.
  • Collaborate with application engineers and developers to identify root causes and implement durable fixes.
  • Incident Management & Root Cause Analysis
  • Participate as a Subject Matter Advisor during production incidents and outages.
  • Provide insights backed by system monitoring, code review, and database analysis.
  • Support post‑mortem reviews and drive follow‑up actions.
  • Automation & Tooling
  • Automate operational tasks such as monitoring, alerts, and recovery processes.
  • Build scripts and internal tools to eliminate manual toil and improve operational efficiency.
  • Monitoring & Observability
  • Implement telemetry and observability practices to track system health, latency, and error rates.
  • Manage the Dynatrace platform and its integrations with application services.
  • Support teams in designing dashboards and visualization setups.
  • Security & Compliance
  • Work with Security teams to ensure systems comply with regulatory and industry standards (e.g., PCI‑DSS, GDPR).
  • Implement necessary access controls, encryption, and audit capabilities within SRE scope.
  • Capacity Planning & Performance Optimization
  • Analyze usage trends to forecast demand and support scaling decisions.
  • Contribute to cost‑performance optimization efforts across infrastructure and applications.
  • Collaborate closely with development, QA, and infrastructure teams to embed reliability into the SDLC.
  • Documentation & Knowledge Sharing
  • Maintain clear and up‑to‑date operational documentation, runbooks, and architecture diagrams.
  • Champion SRE principles across the organization to foster resilience and accountability.
Job Requirements
Education
  • Bachelor’s degree in Computer Science, Software Engineering, IT, or related fields.
Experience
  • 3–5 years of experience in SRE, DevOps, or Software Engineering roles.
  • Experience supporting front‑end applications in production environments, ideally within financial services or other regulated industries.
Technical Skills
  • Strong understanding of front‑end performance monitoring and instrumentation.
  • Hands‑on experience with Real User Monitoring (RUM), Synthetic Monitoring, and APM tools (e.g., Dynatrace, New Relic, Datadog).
  • Proficiency in building dashboards and alerts using Dynatrace, Grafana, Prometheus, Elastic Stack, or Splunk.
  • Familiarity with OpenTelemetry for distributed tracing.
  • Scripting skills in Python, Bash, or JavaScript.
  • Experience with CI/CD pipelines (e.g., GitHub Flow).
  • Practical experience with cloud technologies (AWS or Azure).
  • Knowledge of Docker and Kubernetes.
  • Understanding of secure coding practices for front‑end applications.
  • Awareness of financial compliance standards such as PCI‑DSS.
Why Join Us?
  • Be part of a high‑impact team shaping system resilience across the enterprise.
  • Work with modern observability and automation technologies.
  • Influence engineering culture through SRE best practices.
  • Opportunities to innovate and drive real improvements in system reliability.

Build a career with us as we help our customers and the community live Healthier, Longer, Better Lives. You must provide all requested information, including Personal Data, to be considered for this career opportunity. Failure to provide such information may influence the processing and outcome of your application. You are responsible for ensuring that the information you submit is accurate and up‑to‑date.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Hong Kong and Macau • Kuala Lumpur

On-site
MYR 70,000 - 90,000
SRE Lead
SRE Lead

Chubb Ltd. • Malaysia

On-site
MYR 240,000 - 420,000
SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Career Wise • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior Site Reliability Engineer (Application)
Senior Site Reliability Engineer (Application)

Guidewire Software • Kuala Lumpur

Hybrid
MYR 180,000 - 300,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ryt Bank • Kuala Lumpur

On-site
MYR 120,000 - 180,000
System Reliability Engineer
System Reliability Engineer

Michael Page • Kuala Lumpur

On-site
MYR 54,000 - 120,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Esker • Kuala Lumpur

Hybrid
MYR 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ryt Bank • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Regional Site Reliability Engineer (SRE)
Regional Site Reliability Engineer (SRE)

Zuspresso (M) Sdn Bhd • Shah Alam

On-site
MYR 120,000 - 180,000