Cloud Site Reliability Engineer (SRE)

Singapore Pools

Singapore

On-site

SGD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Total rewards
Health & wellness benefits
Learning opportunities
Volunteerism

Job summary

Singapore Pools (Pte) Ltd is seeking a Site Reliability Engineer to architect scalable cloud infra, manage the DevOps toolchain and drive observability across hybrid environments.

You will implement GitOps, IaC, GenAI-enabled automation and work with AWS/Azure, aiming for maximum uptime, efficient FinOps, and robust disaster recovery with comprehensive runbooks.

Qualifications

  • Degree-qualified in Computer Science, Engineering, Information Science or related IT discipline.
  • 4 to 7 years hands-on experience in software engineering, cloud architecture, DevOps, or SRE roles.
  • Certifications such as ITIL, FinOps Practitioner, AWS/Azure certifications listed in the job ad.
  • Experience building scalable hybrid cloud infrastructure (AWS and Azure).
  • Production software development in Python, Golang or Java; API-based services and serverless on AWS Lambda / Azure Functions.
  • IaC tooling (Terraform, Ansible, CloudFormation) and Kubernetes; agile methodologies.
  • Git with advanced branching and GitOps practices.
  • Observability with Prometheus, Grafana, Dynatrace, Splunk, InfluxDB and time-series databases.
  • Strong communication and stakeholder collaboration.
  • Interest in trends affecting Cloud Engineering and SRE.

Responsibilities

  • Software Engineering & Automation: write code and automate API-based apps with GenAI integrations.
  • Reliability & Operations: implement SRE observability, define SLOs, participate in on-call and postmortems.
  • Hybrid Cloud Infrastructure: enforce rigorous IaC reviews for fault-tolerant apps.
  • CI/CD & Toolchain: maintain DevOps systems, GitOps workflows, and security scans in pipelines.
  • Observability: build dashboards by unifying telemetry from monitoring tools and cloud metrics.
  • FinOps: provide reporting to optimize hybrid cloud spending.
  • Resilience & Documentation: DR, backups, capacity planning, and runbooks.

Skills

Python
Golang
Java
AWS
Azure
Kubernetes
Git
GitOps
IaC
Automation
Observability
GenAI
TimeseriesDB

Education

Degree in CS/Engineering/IS

Tools

Terraform
Ansible
CloudFormation
Prometheus
Grafana
Dynatrace
Splunk
InfluxDB
AWS Lambda
Azure Functions
Docker
Kubernetes

Job description

Company: Singapore Pools (Pte) Ltd

Work that powers communities.

Who We Are

Singapore Pools was established by the Singapore government on 23 May 1968 to provide safe and trusted betting to counter illegal gambling. As a not-for-profit organisation, it makes contributions to the Tote Board to fund a wide range of causes in social service, community development, sports, arts, education and health sectors.

Since 2004, over $5 billion have been channelled to the Tote Board. In addition, Singapore Pools also contributes about $2 billion annually to the Government in the form of taxes and duties. Its responsible gaming practices have been awarded the highest level of certification (Level 4) by the World Lottery Association’s Responsible Gaming Framework since 2012.

Since inception, Singapore Pools’ staff have a long-standing commitment to doing good and giving back to those in need. Staff volunteers support activities held all year round, from helping disadvantaged children, youth-at-risk, underprivileged families, and elderly, to conserving the environment.

Job Purpose

The Site Reliability Engineer (SRE) drives enterprise operational resilience by architecting scalable cloud infrastructure, managing the enterprise DevOps toolchain, and ensuring centralized system observability across hybrid cloud environments. Leveraging a strong software engineering background, the SRE implements GitOps methodologies and develops custom automation applications and API-based services. By heavily utilizing Infrastructure as Code (IaC) and GenAI tools, this role eliminates manual operations, drives efficiency, optimizes cloud expenditures, and ensures maximum system uptime against stringent SLAs.

What You'll Do
  • Software Engineering & Automation: Write code and develop automated, API-based applications (including GenAI integrations) to streamline operational reporting and eliminate recurring manual tasks.
  • Reliability & Operations: Implement SRE best practices for observability, availability, performance, and incident response. Define, measure, and govern Service Level Objectives (SLOs) and Error Budgets in collaboration with product engineering teams. Participate in on-call rotations, execute postmortem/RCAs, and identify/fix production bottlenecks.
  • Hybrid Cloud Infrastructure: Support product teams in building fault-tolerant applications by enforcing infrastructure deployment via rigorous IaC code reviews.
  • CI/CD & Toolchain: Maintain DevOps systems and enforce GitOps workflows and integrate automated security/vulnerability scanning into CI/CD pipelines for all application, container, and IaC deployment approvals.
  • Observability: Create and maintain consolidated operations dashboards by integrating telemetry from disparate monitoring tools and native AWS/Azure metrics.
  • FinOps: Generate actionable FinOps reporting to track and optimize hybrid cloud spending.
  • Resilience & Documentation: Execute disaster recovery (DR), backup, redundancy, and capacity planning strategies while maintaining high-quality runbooks and operational documentation.
Who You Are
  • Degree-qualified in Computer Science, Engineering, Information Science or related IT Discipline, with 4 to 7 years of proven, hands-on experience in software engineering, cloud architecture, DevOps, or SRE roles.
  • Hold professional certifications such as ITIL, FinOps Certified Practitioner, AWS Certified Solutions Architect (Associate or Professional), AWS Certified DevOps Engineer (Professional), AWS Certified CloudOps Engineer (Associate), Microsoft Certified Azure Administrator Associate (AZ-104), Azure Solutions Architect Expert (AZ-305), or DevOps Engineer Expert (AZ-400).
  • Cloud Architecture & Engineering: Deep hands-on experience building scalable hybrid cloud infrastructures (AWS and Azure) and containerization. Strong understanding of modern hosting, networking design patterns, and applying the Six Pillars of operational excellence across environments.
  • Software Engineering: Strong background building production-level software in Python, Golang, or Java. Experience developing and deploying API-based services and serverless applications on AWS Lambda and Azure Functions.
  • Automation & Infrastructure as Code (IaC): High proficiency in IaC and configuration management tools (e.g., Terraform, Ansible, CloudFormation) and container orchestration systems (e.g., Kubernetes). Proven capability in enforcing deployment pipelines through rigorous code review processes and working within Agile methodologies.
  • Version Control: Proficient in Git, including advanced branching strategies and GitOps paradigms.
  • Systems Knowledge: Deep expertise in architecting and managing centralized Observability platforms, utilizing GenAI, time-series databases, and diverse monitoring tools (e.g., Prometheus, Grafana, Dynatrace, Splunk, InfluxDB) to trace distributed cloud applications. Solid understanding of Database Administration and Networking architecture.
  • Professional & Interpersonal Skills: Ability to make sound, logical, data-based decisions on complex issues while considering risks. Strong communication and interpersonal skills to collaborate and build relationships with internal and external stakeholders.
  • Strong interest in technological trends and disruptions impacting Cloud Engineering and SRE.
What We Offer
  • Comprehensive total rewards package
  • Health & wellness benefits
  • Continuous learning and upskilling opportunities
  • Volunteerism and community initiatives

Only shortlisted candidates will be contacted for further career conversations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
Assistant Manager, OpsWatch Unit
Assistant Manager, OpsWatch Unit

Singapore Pools • Singapore

On-site
SGD 120,000 - 180,000
Total rewards
Health benefits
Learning opportunities
+1
Cloud SRE & Automation Engineer
Cloud SRE & Automation Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
Senior Engineer, Retail and Customer Interaction Delivery
Senior Engineer, Retail and Customer Interaction Delivery

Singapore Pools • Singapore

On-site
SGD 120,000 - 180,000
Comprehensive total rewards package
Health & wellness benefits
Continuous learning and upskilling
+1
Platform Engineer (Cloud SRE Ops)
Platform Engineer (Cloud SRE Ops)

Assurity Trusted Solutions Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Annual Leave
Family Care Leave
Birthday Leave
+1
Senior Site Reliability Engineer / SRE Lead
Senior Site Reliability Engineer / SRE Lead

Reolink Technology Pte. Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Insurance Coverage
Yearly Bonus & Performance Bonus
Engineer, Retail and Customer Interaction Delivery
Engineer, Retail and Customer Interaction Delivery

Singapore Pools • Singapore

On-site
SGD 60,000 - 100,000
Comprehensive total rewards package
Health & wellness benefits
Continuous learning and upskilling
+1
Senior Site Reliability Engineer / SRE Leader
Senior Site Reliability Engineer / SRE Leader

REOLINK TECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Insurance Coverage
Yearly Bonus
Performance Bonus
Platform Engineer (Cloud SRE Ops)
Platform Engineer (Cloud SRE Ops)

Assurity Trusted Solutions • Singapore

Hybrid
SGD 120,000 - 180,000
Annual leave with family care
Birthday leave
Learning culture
Site Reliability Engineer
Site Reliability Engineer

TEKsystems • Singapore

Hybrid
SGD 120,000 - 180,000