An application made for this job — a tailored resume and cover letter that speak straight to the posting.
SGS (Malaysia) Sdn Bhd is seeking a Site Reliability Engineer to design, build, and maintain scalable, cloud-native systems. You will drive automation, observability, and incident response in a fast-paced environment, collaborating with software and IT teams to ensure availability and resilience of mission-critical services.
You will own CI/CD pipelines and IaC, manage Kubernetes clusters, and implement security and disaster recovery practices.
We are seeking a highly skilled and proactive Site Reliability Engineer (SRE) to join our technology team. In this role, you will be instrumental in designing, building, and maintaining scalable, reliable, and secure cloud-portable systems. You will collaborate closely with software engineering and IT teams to ensure the availability, performance, and resilience of our mission-critical services. This is an exciting opportunity to champion best practices in automation, observability, and incident response within a dynamic, fast-paced, cloud-native environment.
Design, build, and maintain highly available, scalable systems and infrastructure.
Develop and implement automation for deployment, monitoring, management, and alerting.
Partner with development teams to ensure reliability and operational excellence throughout the application lifecycle.
Monitor system performance, proactively identify issues, and drive root cause analysis and resolution of incidents.
Define and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
Create and maintain robust documentation for systems, processes, and incident reports.
Manage and maintain CI/CD pipelines and infrastructure as code (IaC).
Champion best practices for security, compliance, and disaster recovery.
Participate in on-call rotations and respond to production incidents.
Continuously seek opportunities to improve system reliability, scalability, and efficiency.
Candidates should already possess the required work eligibility, such as Sarawakian status.
3+ years of experience in a Site Reliability or DevOps role.
Strong expertise in Linux/Unix system administration and TCP/IP networking.
Hands-on experience with at least one public cloud provider (AWS, Azure, or GCP).
Proven experience managing Kubernetes (K8s) clusters, including upgrades, scaling, and troubleshooting.
Experience with containerization technologies (Docker, Padman, etc.).
Proficiency with scripting or programming languages (Go, java, Python, Bash, etc.).
Solid understanding of monitoring, logging, and observability tools (Prometheus, Grafana, ELK, Datadog, etc.).
Experience implementing and managing GitOps workflows and automation tools (e.g., FluxCD, ArgoCD).
Familiarity with configuration management/IaC tools (Terraform, Ansible, Helm, etc.).
Excellent troubleshooting and problem-solving skills.
Hirer responsiveness Salary match Number of applicants
Your application will include the following questions:
Engineering & Project Management Services 1,001-5,000 employees
SGS is the world’s leading Testing, Inspection and Certification company.
We operate a network of over 2,500 laboratories and business facilities across 115 countries, supported by a team of over 100,000 dedicated professionals.
With more than 145 years of service excellence, we combine the precision and accuracy that define Swiss companies to help organizations achieve the highest standards of quality, compliance and sustainability.
Our brand promise – when you need to be sure – underscores our commitment to trust, integrity and reliability, enabling businesses to thrive with confidence.
SGS is the world’s leading Testing, Inspection and Certification company.
We operate a network of over 2,500 laboratories and business facilities across 115 countries, supported by a team of over 100,000 dedicated professionals.
With more than 145 years of service excellence, we combine the precision and accuracy that define Swiss companies to help organizations achieve the highest standards of quality, compliance and sustainability.
Our brand promise – when you need to be sure – underscores our commitment to trust, integrity and reliability, enabling businesses to thrive with confidence.
What can I earn as a Site Reliability Engineer