Site Reliability Engineer (SRE)

SGS (Malaysia) Sdn Bhd

Kuching

On-site

MYR 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SGS (Malaysia) Sdn Bhd is seeking a Site Reliability Engineer to design, build, and maintain scalable, cloud-native systems. You will drive automation, observability, and incident response in a fast-paced environment, collaborating with software and IT teams to ensure availability and resilience of mission-critical services.

You will own CI/CD pipelines and IaC, manage Kubernetes clusters, and implement security and disaster recovery practices.

Qualifications

  • 3+ years in a Site Reliability or DevOps role
  • Strong Linux/Unix system administration and TCP/IP networking
  • Hands-on with at least one public cloud provider (AWS, Azure, or GCP)
  • Proven experience managing Kubernetes clusters, including upgrades, scaling, and troubleshooting
  • Experience with monitoring, logging, and observability tools (Prometheus, Grafana, ELK, Datadog)
  • Experience implementing and managing GitOps workflows and automation tools (FluxCD, ArgoCD)
  • Familiarity with configuration management/IaC tools (Terraform, Ansible, Helm)
  • Excellent troubleshooting and problem-solving skills

Responsibilities

  • Design, build, and maintain highly available, scalable systems and infrastructure
  • Develop and implement automation for deployment, monitoring, management, and alerting
  • Partner with development teams to ensure reliability and operational excellence
  • Monitor system performance, identify issues, and drive root cause analysis and resolution
  • Define and track SLIs, SLOs, and SLAs
  • Create and maintain robust documentation for systems, processes, and incident reports
  • Manage CI/CD pipelines and infrastructure as code (IaC)
  • Champion security, compliance, and disaster recovery practices
  • Participate in on-call rotations and respond to production incidents
  • Continuously seek opportunities to improve system reliability and efficiency

Skills

Linux/Unix administration
TCP/IP networking
Scripting (Go/Java/Python/Bash)
Monitoring/Observability
GitOps workflows

Tools

Kubernetes
Docker
Prometheus
Grafana
ElK/Datadog
FluxCD
ArgoCD
Terraform
Ansible
Helm

Job description

We are seeking a highly skilled and proactive Site Reliability Engineer (SRE) to join our technology team. In this role, you will be instrumental in designing, building, and maintaining scalable, reliable, and secure cloud-portable systems. You will collaborate closely with software engineering and IT teams to ensure the availability, performance, and resilience of our mission-critical services. This is an exciting opportunity to champion best practices in automation, observability, and incident response within a dynamic, fast-paced, cloud-native environment.

Key responsibilities

Design, build, and maintain highly available, scalable systems and infrastructure.

Develop and implement automation for deployment, monitoring, management, and alerting.

Partner with development teams to ensure reliability and operational excellence throughout the application lifecycle.

Monitor system performance, proactively identify issues, and drive root cause analysis and resolution of incidents.

Define and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).

Create and maintain robust documentation for systems, processes, and incident reports.

Manage and maintain CI/CD pipelines and infrastructure as code (IaC).

Champion best practices for security, compliance, and disaster recovery.

Participate in on-call rotations and respond to production incidents.

Continuously seek opportunities to improve system reliability, scalability, and efficiency.

About you

Candidates should already possess the required work eligibility, such as Sarawakian status.

3+ years of experience in a Site Reliability or DevOps role.

Strong expertise in Linux/Unix system administration and TCP/IP networking.

Hands-on experience with at least one public cloud provider (AWS, Azure, or GCP).

Proven experience managing Kubernetes (K8s) clusters, including upgrades, scaling, and troubleshooting.

Experience with containerization technologies (Docker, Padman, etc.).

Proficiency with scripting or programming languages (Go, java, Python, Bash, etc.).

Solid understanding of monitoring, logging, and observability tools (Prometheus, Grafana, ELK, Datadog, etc.).

Experience implementing and managing GitOps workflows and automation tools (e.g., FluxCD, ArgoCD).

Familiarity with configuration management/IaC tools (Terraform, Ansible, Helm, etc.).

Excellent troubleshooting and problem-solving skills.

Unlock job insights

Hirer responsiveness Salary match Number of applicants

Your application will include the following questions:

  • Which of the following statements best describes your right to work in Malaysia?
  • What's your expected monthly basic salary?
  • How many years' experience do you have as a Site Reliability Engineer?
  • Which of the following types of qualifications do you have?
  • Do you have Root Cause Analysis experience?
  • Do you have CI/CD experience?
  • Do you have experience using Linux Operating Systems?

Engineering & Project Management Services 1,001-5,000 employees

SGS is the world’s leading Testing, Inspection and Certification company.

We operate a network of over 2,500 laboratories and business facilities across 115 countries, supported by a team of over 100,000 dedicated professionals.

With more than 145 years of service excellence, we combine the precision and accuracy that define Swiss companies to help organizations achieve the highest standards of quality, compliance and sustainability.

Our brand promise – when you need to be sure – underscores our commitment to trust, integrity and reliability, enabling businesses to thrive with confidence.

SGS is the world’s leading Testing, Inspection and Certification company.

We operate a network of over 2,500 laboratories and business facilities across 115 countries, supported by a team of over 100,000 dedicated professionals.

With more than 145 years of service excellence, we combine the precision and accuracy that define Swiss companies to help organizations achieve the highest standards of quality, compliance and sustainability.

Our brand promise – when you need to be sure – underscores our commitment to trust, integrity and reliability, enabling businesses to thrive with confidence.

What can I earn as a Site Reliability Engineer

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

LAVU TECH SOLUTIONS SDN. BHD. • Petaling Jaya

On-site
MYR 180,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

Lavu Tech Solutions • Selangor

On-site
MYR 180,000 - 280,000
Site Reliability Engineer
Site Reliability Engineer

Aisling Group • Kuala Lumpur

On-site
Senior Site Reliability Engineer - Scale-Up Platform
Senior Site Reliability Engineer - Scale-Up Platform

Aisling Group • Kuala Lumpur

On-site
Cloud-Native Site Reliability Engineer
Cloud-Native Site Reliability Engineer

SGS (Malaysia) Sdn Bhd • Kuching

On-site
MYR 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Esker • Kuala Lumpur

Hybrid
MYR 120,000 - 180,000
Site Reliability Engineer (SRE) Esker Asia · Kuala Lumpur, Malaysia ·
Site Reliability Engineer (SRE) Esker Asia · Kuala Lumpur, Malaysia ·

Esker, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Site Reliability Engineer: Build Resilient, Scalable Systems
Site Reliability Engineer: Build Resilient, Scalable Systems

Career Wise • Kuala Lumpur

On-site
MYR 80,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Experian Marketing Services (Malaysia) Sdn Bhd • Cyberjaya

Hybrid
MYR 180,000 - 300,000
Hybrid work model
Discretionary bonus
Great compensation package
+1
Site Reliability Engineer: Build Reliable, Scalable Systems
Site Reliability Engineer: Build Reliable, Scalable Systems

Setel Ventures • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Leisure area with video games
Casual dress (jeans)
Pantry with coffee, tea and snacks
+2