Enable job alerts via email!

Site Reliability Engineer (SRE) | KL

Hunters International Sdn Bhd

Kuala Lumpur

On-site

MYR 200,000 - 250,000

Full time

9 days ago

Generate a tailored resume in minutes

Land an interview and earn more. Learn more

Job summary

A leading recruitment agency in Kuala Lumpur is seeking a Site Reliability Engineer (SRE) to ensure the reliability and performance of critical services. The ideal candidate must be proficient in Mandarin and programming languages like Python, Golang, and Java. This role emphasizes key SRE practices such as SLOs and SLIs, alongside strong problem-solving skills. The position offers a competitive salary of up to MYR 19,000 based on relevant experience.

Qualifications

Proficiency in Mandarin is a must for liaison.
Experience in system architecture and design.
Strong understanding of SRE principles.

Responsibilities

Design and implement resilient system architectures.
Develop automation tools to enhance operational efficiency.
Conduct thorough post-mortem analyses following incidents.
Collaborate with teams to establish best practices in incident management.

Skills

Mandarin

Python

Golang

Java

Linux system administration

Networking concepts

Scripting for automation

Cloud platforms

DevOps practices

About the job Site Reliability Engineer (SRE) | KL

Overview:

As a Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of
continuous learning and accountability.

Responsibilities

Design and implement resilient system architectures that support high availability and scalability.
Develop automation tools and scripts to enhance operationalefficiency and reduce manual effort.
Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
Conduct thorough post-mortem analyses following incidents, driving continuousimprovement through root cause identification and solution implementation.
Collaborate with development and operations teams to establish best practices in system reliability and incident management.
Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs),maintaining high standards of service delivery.
Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.

Requirements:

Proficiency in Mandarin is a must, in order to liaise with stakeholders from China.
Proficiency in programming languages such asPython, Golang, Java,or similar, focusing on operational efficiency.
Demonstrated experience insystem architecture and design,prioritizing reliability, and scalability.
Strong understanding ofSRE principles, includingSLOs, SLIs, toil reduction, and incident post-mortems.
Experience withcloud environments (e.g., AWS, Azure, Google Cloud)and their operational management.
Strong expertise inLinuxsystem administration.
Proven experience introubleshooting application support issueswith a focus on performance and connectivity.
Familiarity withnetworking conceptsand effective troubleshooting techniques.
Excellent problem-solving abilities and a proactive approach to operational challenges.
Ability to work independently while effectively collaborating within a team environment.
Familiarity with monitoring tools and performance optimization techniques.
Experience inscripting or automation for system administration tasks.
Knowledge ofnetworking concepts and troubleshooting methodologies.
Hands-on knowledge ofcloud platforms (e.g., AWS, Azure, Google Cloud)and their services.
Familiarity withDevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.

Location:

To be based in client's site (Klang Valley)

Remuneration:

Up to MYR 19,000 (Based on relevant experience)

Consultant in Charge

May Chong | may.chong@hunters-in.com | 012 280 1717

This is a contract position with the possibility to be absorbed as a permanent staff.

Get your free, confidential resume review.

or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.