Platform Engineer (Cloud SRE Ops)

Assurity Trusted Solutions

Singapore

Hybrid

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Annual leave with family care
Birthday leave
Learning culture

Job summary

Assurity Trusted Solutions seeks a seasoned Infrastructure Engineer/SRE to design, build, and operate large-scale government production systems in a demanding environment. You will lead incident response, automation, and reliability initiatives across cross-functional teams, ensuring high availability and efficient performance.

The role emphasizes on-call duties, proactive monitoring, and collaboration with software engineers and DevOps to prevent issues and accelerate recovery.

Qualifications

  • 6+ years of experience in technology operations as an Infrastructure Engineer or Site Reliability Engineer with large-scale production systems.
  • Expertise in building and operating automated monitoring and incident detection systems, runbooks and incident management processes.
  • Experience designing automation solutions using provisioning tools, CI/CD, and scripting languages.
  • Experience leading highly complex technical projects with multiple dependencies and stakeholders.
  • Knowledge of Agile development environments and delivering high-quality services quickly.
  • Proficiency in building and managing highly available and scalable IT infrastructure and/or applications, with containerization and virtualization.
  • Proficiency in Python, PowerShell, or Ruby.
  • Experience with Infrastructure as Code tools such as SaltStack, Puppet, Terraform, or Ansible.

Responsibilities

  • Build Service Level Indicators (SLI), Service Level Objectives (SLO), error budgets, and post-mortem processes for incidents.
  • Provide on-call reliability and performance support for critical Government services.
  • Analyze OS and application metrics/logs for capacity planning and performance tuning.
  • Create automation to manage services, infrastructure, and applications.
  • Improve reliability through proactive monitoring and issue prevention measures.
  • Drive continuous improvement in SRE practices and system performance.
  • Develop SRE playbooks for the Whole-of-Government reference.
  • Work with cross-functional teams of software and infrastructure engineers and DevOps.

Skills

DevOps
Infrastructure engineering
SRE
Python
PowerShell
Ruby
CI/CD
Monitoring
Automation
Incident response

Tools

SaltStack
Puppet
Terraform
Ansible

Job description

In Digital Resiliency Engineering (DRE), we combine software and systems engineering to build and operate large-scale and distributed systems designed and/or built by the Singapore Government. We ensure Government services are reliable, meets expected performance and satisfy customer needs.

If you are someone with strong DevOps, Infrastructure engineering and/or SRE background, have experience operating mission critical production technology infrastructure at scale, and are looking for opportunities to work with a team of practitioners and leading industry experts, we welcome you to join us.

In this role, you will build central services for observability and automation of infrastructure services. You will be part of a rotation with other engineers in providing rapid response to major incidents impacting critical Government Services. You will provide technical leadership for the team and work closely with technical leads to operate highly available solutions. You will also provide guidance to other team member on managing availability and performance of mission critical services, building automation and monitoring solutions to prevent problem recurrence, and building automated responses for non-exceptional service conditions.

You will also manage execution of project priorities, deadlines and deliverables. You will also lead designs of major components, systems and features to improve availability, scalability, latency and efficiency of services design and built by the Government.

Key Responsibilities:

  • Build Service Level Indicators (SLI), Service Level Objective (SLO), Error Budgets, and Post-mortem Incident processes.
  • As part of an on-call roster, ensure reliability and performance of critical Government Services. Provide operational support and engineering for large-scale and distributed systems to drive incidents resolution effectively.
  • Gather and analyse metrics and logs from Operating Systems and/or applications for capacity planning, performance tuning and fault isolation.
  • Build automation to manage services, infrastructure, and/or applications.
  • Improve reliability and quality of services using proactive monitoring.
  • Measure and optimize system performance, with continuous improvement and pushing SRE practice forward.
  • Build SRE playbook for the Whole-of-Government to leverage as reference for SRE.
  • Identify potential and emerging technologies relevant to innovation for the Government.
  • Work in a cross-functional service team consisting of software engineers, infrastructure engineers, DevOps, and other specialists.
  • 6+ years of experience in technology operations as an Infrastructure Engineer or Site Reliability Engineer - with experience operating large-scale mission critical production systems.
  • Expertise in building and operating automated monitoring and incident detection systems, creating runbooks and running incident management processes.
  • Expertise in designing automation solutions using provisioning tools, continuous integration tools (CI/CD), and scripting languages.
  • Experience leading highly complex technical projects with multiple dependencies and stakeholders
  • Knowledgeable and experienced in working within an Agile development environment, focusing on dynamic and rapid quality delivery.
  • Proficient in building and managing highly available and scalable IT infrastructure and/or application, with knowledge in Container and Virtualization technologies.
  • Proficiency in Python, PowerShell, or Ruby.
  • Proficiency with Infrastructure as Code (IaC) tools such as SaltStack, Puppet, Terraform, or Ansible.
  • Able to work independently and deliver results within specified deadlines.
  • Ability to prioritize work and strong problem-solving skills.
  • Good to have communicate skills, both verbally and in writing to users, vendors and management.
  • Ability to communicate complex interaction concepts clearly and persuasively across different audience and varies levels in GovTech.

Join us and discover a meaningful and exciting career with Assurity Trusted Solutions!

The remuneration package will commensurate with your qualifications and experience.

We thank you for your interest and please note that only shortlisted candidates will be notified.

  • A wholly-owned subsidiary of GovTech.
  • We promote a learning culture and encourage you to grow and learn.
  • Annual Leave Benefits with additional perks such as Family Care and Birthday Leave.
  • Contract Staff enjoys the same benefits as Permanent Employees.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer (Cloud SRE Ops)
Platform Engineer (Cloud SRE Ops)

Assurity Trusted Solutions Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Annual Leave
Family Care Leave
Birthday Leave
+1
Platform Engineer (SRE/DevOps)
Platform Engineer (SRE/DevOps)

Assurity Trusted Solutions Pte Ltd • Singapore

On-site
SGD 90,000 - 140,000
Annual Leave
Family Care
Birthday Leave
Platform Engineer (SRE/DevOps)
Platform Engineer (SRE/DevOps)

Assurity Trusted Solutions • Singapore

On-site
SGD 90,000 - 150,000
Annual Leave Benefits with Family Care
Birthday Leave
Contract Staff benefits equal to perm
Cloud Site Reliability Engineer
Cloud Site Reliability Engineer

SINGAPORE POOLS (PRIVATE) LIMITED. • Singapore

On-site
SGD 120,000 - 160,000
Total rewards
Health benefits
Learning opportunities
+1
DevSecOps Engineer - A26175
DevSecOps Engineer - A26175

Activate Interactive • Singapore

On-site
SGD 60,000 - 90,000
Fun working environment
Employee Wellness Program
Structured development framework
Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

Singapore Pools • Singapore

On-site
SGD 120,000 - 180,000
Total rewards
Health & wellness benefits
Learning opportunities
+1
Senior DevOps Engineer: Cloud Infra & CI/CD
Senior DevOps Engineer: Cloud Infra & CI/CD

GovTech Singapore • Singapore

On-site
SGD 80,000 - 120,000
Flexible work arrangements
Employee wellness programs
Comprehensive leave benefits
Platform & SRE Engineer (Database and Search Platform)
Platform & SRE Engineer (Database and Search Platform)

Csit • Singapore

On-site
SGD 90,000 - 130,000
SRE Engineer
SRE Engineer

BOUNTEOUSXACCOLITE SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 180,000
Infra Security Engineer
Infra Security Engineer

Centre for Strategic Infocomm Technologies • Singapore

On-site
SGD 75,000 - 110,000