Site Reliability Engineer (SRE)

Career Wise

Kuala Lumpur

On-site

MYR 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology firm in Kuala Lumpur is seeking a skilled Site Reliability Engineer (SRE) to ensure the reliability and performance of critical services. The ideal candidate should have a minimum of 3 years of experience in system architecture, strong programming skills in languages such as Python and Golang, and a solid understanding of SRE principles including SLOs and SLIs. The role involves collaboration with diverse teams, enhancement of operational efficiency through automation, and contributing to a culture of continuous learning and improvement.

Qualifications

  • At least 3 years of experience in a related field.
  • Strong understanding of SRE principles like SLOs and SLIs.
  • Proficiency in programming languages focusing on operational efficiency.

Responsibilities

  • Design and implement resilient system architectures for high availability.
  • Develop automation tools to enhance operational efficiency.
  • Conduct post-mortem analyses following incidents for continuous improvement.

Skills

Python
Golang
Java
Linux system administration
Troubleshooting
Cloud environments (AWS, Azure, Google Cloud)

Job description

About the job Site Reliability Engineer (SRE)

As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability.

Key Responsibilities:

  • Design and implement resilient system architectures that support high availability and scalability.
  • Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
  • Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
  • Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation.
  • Collaborate with development and operations teams to establish best practices in system reliability and incident management.
  • Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
  • Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery.
  • Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
  • Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.

Qualifications:

  • Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency.
  • Minimum experience of 3 years and above in related field.
  • Demonstrated experience in system architecture and design, prioritizing reliability, and scalability.
  • Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems.
  • Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management.
  • Strong expertise in Linux system administration.
  • Proven experience in troubleshooting application support issues with a focus on performance and connectivity.
  • Familiarity with networking concepts and effective troubleshooting techniques.
  • Excellent problem-solving abilities and a proactive approach to operational challenges.
  • Ability to work independently while effectively collaborating within a team environment.

Preferred Skills:

  • Familiarity with monitoring tools and performance optimization techniques.
  • Experience in scripting or automation for system administration tasks.
  • Knowledge of networking concepts and troubleshooting methodologies.
  • Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services.
  • Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
SRE Lead
SRE Lead

Chubb Ltd. • Malaysia

On-site
MYR 240,000 - 420,000
Site Reliability Engineer: Build Resilient, Scalable Systems
Site Reliability Engineer: Build Resilient, Scalable Systems

Career Wise • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AirAsia rewards • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Regional Site Reliability Engineer (SRE)
Regional Site Reliability Engineer (SRE)

Zuspresso (M) Sdn Bhd • Shah Alam

On-site
MYR 120,000 - 180,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
Site (Software) Reliability Engineer (SRE)
Site (Software) Reliability Engineer (SRE)

Realtek • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Ryt Bank • Kuala Lumpur

On-site
MYR 120,000 - 180,000
SRE Engineer (DevOps)
SRE Engineer (DevOps)

Ant International • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Engineering Manager – Platform & SRE
Engineering Manager – Platform & SRE

INSCALE • Kuala Lumpur

On-site
MYR 350,000 - 650,000