Site Reliability Engineer (6 month contarcts)

Zühlke Group

Singapore

On-site

SGD 90,000 - 130,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Zühlke Group in Singapore seeks an experienced Platform Reliability Engineer to monitor platforms, troubleshoot incidents, perform root-cause analyses, and drive reliability improvements. You will automate runbooks and participate in on-call rotations.

The role requires strong cloud and container skills, IaC with Terraform/CloudFormation, CI/CD with GitHub Actions, and experience with CloudWatch, Splunk and Datadog; you will collaborate with agile teams to deliver dependable services.

Qualifications

  • Possess a degree in Computer Science/Information Technology or related fields.
  • At least 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
  • At least 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous).
  • Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
  • Proficiency with CI/CD workflows and GitHub Actions.
  • Knowledge of artifact repository management systems such as Frog.
  • Strong Linux administration skills and Shell scripting expertise.
  • Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
  • Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root-cause analysis, and post-incident follow-up.
  • Experience in diagnosing and resolving complex system issues across multiple technology layers.
  • Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
  • Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
  • Willingness to participate in an on-call support rotation for critical production services, where required.
  • Able to optimize developer workflows and enhance developer experience.
  • Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
  • Excellent communication skills to work effectively with diverse engineering teams.
  • Strong team-player mindset, focused on leveraging experience to help the team succeed.
  • Possess positive learning and collaborative mindset.
  • Strong analytical, problem-solving and troubleshooting skills.
  • Good written and verbal communication skills.
  • Agile, fast learner and able to adapt to changes.

Responsibilities

  • Responsibilities include platform monitoring, incident troubleshooting and resolution, root-cause analysis, service reliability improvements, operational automation, runbook maintenance, and on-call support where required.
  • The role will help maintain platform availability, stability, and performance while reducing manual operational effort.

Skills

AWS
Docker
Kubernetes
Terraform
GitHub Actions
Linux
CI/CD

Education

Computer Science/Information Technology degree

Tools

CloudWatch
Datadog
Splunk
Helm
EKS
CloudFormation
Frog
Terraform

Job description

Founded in Switzerland in 1968, Zühlke is a team of colleagues across Europe and Asia, empowering ideas and creating new business models by developing services and products based on new technologies. While we work with the latest technologies on complex business challenges globally, our priority is to nurture what sets us apart: our people.

Working with us, you’ll be part of agile, collaborative teams with the opportunity to deliver transformational impact through technology, engineering excellence, and meaningful client outcomes.

The role
  • Responsibilities include platform monitoring, incident troubleshooting and resolution, root-cause analysis, service reliability improvements, operational automation, runbook maintenance, and on-call support where required.
  • The role will help maintain platform availability, stability, and performance while reducing manual operational effort.
What’s important to us
  • Possess a degree in Computer Science/Information Technology or related fields.
  • At least 3 years of AWS experience with a solid understanding of cloud services and infrastructure management (AWS certifications are advantageous).
  • At least 3 years of experience with containerization technologies such as Docker, Kubernetes, EKS, and Helm (relevant certifications are advantageous.
  • Proven experience with infrastructure as code tools such as Terraform and CloudFormation.
  • Proficiency with CI/CD workflows and GitHub Actions.
  • Knowledge of artifact repository management systems such as Frog.
  • Strong Linux administration skills and Shell scripting expertise.
  • Experience with log aggregation and observability tools such as CloudWatch, Splunk, and Datadog.
  • Working knowledge of service monitoring, alert management, SLIs/SLOs, incident response, root-cause analysis, and post-incident follow-up.
  • Experience in diagnosing and resolving complex system issues across multiple technology layers.
  • Able to troubleshoot production incidents, coordinate timely resolution, and communicate status clearly to technical and business stakeholders.
  • Experience in automating repetitive operational tasks, improving runbooks, and reducing manual support effort.
  • Willingness to participate in an on-call support rotation for critical production services, where required.
  • Able to optimize developer workflows and enhance developer experience.
  • Passion for advocating and implementing best practices in Software Engineering, SRE, and DevOps.
  • Excellent communication skills to work effectively with diverse engineering teams.
  • Strong team-player mindset, focused on leveraging experience to help the team succeed.
  • Possess positive learning and collaborative mindset.
  • Strong analytical, problem-solving and troubleshooting skills.
  • Good written and verbal communication skills.
  • Agile, fast learner and able to adapt to changes.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DevOps / DevSecOps Engineer (6 months contract)
DevOps / DevSecOps Engineer (6 months contract)

Zühlke Group • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer — Cloud, Kubernetes & Automation
Site Reliability Engineer — Cloud, Kubernetes & Automation

Zühlke Group • Singapore

On-site
SGD 90,000 - 130,000
Data Engineer (6 months contract)
Data Engineer (6 months contract)

Zühlke Group • Singapore

On-site
SGD 90,000 - 150,000
Lead DevOps Architect
Lead DevOps Architect

zuhlke engineering pte. ltd. • Singapore

On-site
SGD 100,000 - 180,000
Senior Software Engineer
Senior Software Engineer

zuhlke engineering pte. ltd. • Singapore

On-site
SGD 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Momcozy • Singapore

On-site
SGD 120,000 - 180,000
Competitive compensation
Site Reliability Engineer
Site Reliability Engineer

RemotePeople • Singapore

On-site
SGD 120,000 - 190,000
Senior Field Service Engineer
Senior Field Service Engineer

Swisslog • Singapore

On-site
SGD 54,000 - 78,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Momcozy • Singapore

On-site
SGD 120,000 - 180,000