Platform Engineer (Cloud SRE Ops)

Assurity Trusted Solutions Pte Ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual Leave
Family Care Leave
Birthday Leave
Learning culture

Job summary

GovTech is seeking a Senior Infrastructure/SRE engineer to design, operate and automate large-scale government services. You will lead on-call reliability, implement monitoring, and drive incident response across critical systems.

The role emphasizes automation, scalable architectures and collaboration with software and infrastructure teams. The ideal candidate brings strong scripting skills (Python/PowerShell/Ruby), IaC tooling mastery (SaltStack, Puppet, Terraform, Ansible) and experience in

Qualifications

  • 6+ years in technology operations as Infrastructure/SRE
  • Experience building automated monitoring and incident detection systems
  • Experience with automation solutions and scripting languages
  • Leadership on complex projects in Agile environments
  • Strong communication across GovTech and stakeholders

Responsibilities

  • Build SLIs, SLOs, error budgets and post-mortem processes
  • On-call reliability and performance support for critical services
  • Analyze metrics/logs for capacity planning and performance tuning
  • Develop automation for services, infrastructure and apps
  • Improve reliability with proactive monitoring and incident handling
  • Lead designs to improve availability, latency and efficiency
  • Create SRE playbooks for government-wide use
  • Identify relevant emerging technologies for government use
  • Collaborate with cross-functional teams of engineers and DevOps specialists

Skills

SRE
Automation
CI/CD
Python
PowerShell
Ruby
Terraform
Ansible
SaltStack
Monitoring

Tools

SaltStack
Puppet
Terraform
Ansible

Job description

In Digital Resiliency Engineering (DRE), we combine software and systems engineering to build and operate large-scale and distributed systems designed and/or built by the Singapore Government. We ensure Government services are reliable, meet expected performance and satisfy customer needs.

If you are someone with strong DevOps, Infrastructure engineering and/or SRE background, have experience operating mission critical production technology infrastructure at scale, and are looking for opportunities to work with a team of practitioners and leading industry experts, we welcome you to join us.

In this role, you will build central services for observability and automation of infrastructure services. You will be part of a rotation with other engineers in providing rapid response to major incidents impacting critical Government Services. You will provide technical leadership for the team and work closely with technical leads to operate highly available solutions. You will also provide guidance to other team members on managing availability and performance of mission critical services, building automation and monitoring solutions to prevent problem recurrence, and building automated responses for non-exceptional service conditions. You will also manage execution of project priorities, deadlines and deliverables, and lead designs of major components, systems and features to improve availability, scalability, latency and efficiency of services design and built by the Government.

Key Responsibilities
  • Build Service Level Indicators (SLI), Service Level Objective (SLO), Error Budgets, and Post-mortem Incident processes.
  • As part of an on-call roster, ensure reliability and performance of critical Government Services. Provide operational support and engineering for large-scale and distributed systems to drive incidents resolution effectively.
  • Gather and analyse metrics and logs from Operating Systems and/or applications for capacity planning, performance tuning and fault isolation.
  • Build automation to manage services, infrastructure, and/or applications.
  • Improve reliability and quality of services using proactive monitoring.
  • Measure and optimize system performance, with continuous improvement and pushing SRE practice forward.
  • Build SRE playbook for the Whole-of-Government to leverage as reference for SRE.
  • Identify potential and emerging technologies relevant to innovation for the Government.
  • Work in a cross-functional service team consisting of software engineers, infrastructure engineers, DevOps, and other specialists.
Requirements
  • 6+ years of experience in technology operations as an Infrastructure Engineer or Site Reliability Engineer - with experience operating large-scale mission critical production systems.
  • Expertise in building and operating automated monitoring and incident detection systems, creating runbooks and running incident management processes.
  • Expertise in designing automation solutions using provisioning tools, continuous integration tools (CI/CD), and scripting languages.
  • Experience leading highly complex technical projects with multiple dependencies and stakeholders.
  • Knowledgeable and experienced in working within an Agile development environment, focusing on dynamic and rapid quality delivery.
  • Proficient in building and managing highly available and scalable IT infrastructure and/or application, with knowledge in Container and Virtualization technologies.
  • Proficiency in Python, PowerShell, or Ruby.
  • Proficiency with Infrastructure as Code (IaC) tools such as SaltStack, Puppet, Terraform, or Ansible.
  • Able to work independently and deliver results within specified deadlines.
  • Ability to prioritize work and strong problem-solving skills.
  • Good communication skills, both verbally and in writing, to users, vendors and management.
  • Ability to communicate complex interaction concepts clearly and persuasively across different audience and various levels in GovTech.
Benefits
  • A wholly‑owned subsidiary of GovTech.
  • We promote a learning culture and encourage you to grow and learn.
  • Annual Leave Benefits with additional perks such as Family Care and Birthday Leave.
  • Contract Staff enjoys the same benefits as Permanent Employees.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer (SRE)
Cloud Site Reliability Engineer (SRE)

Singapore Pools • Singapore

On-site
SGD 120,000 - 170,000
Total rewards package
Health & wellness benefits
Continuous learning and upskilling
+1
Senior Platform Infrastructure Engineer, NEA
Senior Platform Infrastructure Engineer, NEA

GovTech Singapore • Singapore

On-site
SGD 90,000 - 130,000
Flexible work arrangements
Leave benefits
Wellness programs
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Autodesk • Singapore

On-site
SGD 80,000 - 120,000
Cloud Infra Engineer
Cloud Infra Engineer

Centre for Strategic Infocomm Technologies (CSIT) • Singapore

On-site
SGD 60,000 - 80,000
Site Reliability Engineer, Engineering Infra - AZ SRE (2027 Graduate)
Site Reliability Engineer, Engineering Infra - AZ SRE (2027 Graduate)

United States Digital Space LLC • Singapore

On-site
SGD 42,000 - 72,000
DevSecOps Engineer - A26187
DevSecOps Engineer - A26187

Activate Interactive • Singapore

On-site
SGD 70,000 - 90,000
Fun working environment
Employee Wellness Program
Growth opportunities
+1
Platform Operations Engineer
Platform Operations Engineer

SERVICE CONNECTIONS HR CONSULTANCY PTE. LTD. • Singapore

On-site
SGD 89,000 - 112,000
Insurance
Operations Lead - SRE, Incident Management
Operations Lead - SRE, Incident Management

Sciente Consulting • Singapore

On-site
SGD 120,000 - 160,000
Software Engineer, Elections Department Singapore
Software Engineer, Elections Department Singapore

GovTech Singapore • Singapore

On-site
SGD 110,000 - 170,000
Flexible work arrangements
Leave benefits and wellness programs
Inclusive workplace culture
Site Reliability Engineer
Site Reliability Engineer

TEKsystems • Singapore

Hybrid
SGD 120,000 - 180,000