Software & Applications Manager (Technical Lead/Supervisor)

optimum solutions (singapore) pte ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

optimum solutions (singapore) pte ltd is seeking an Observability Engineer to lead the enterprise observability and service reliability strategy across infrastructure and application domains. You will drive proactive monitoring, automation, and resilience initiatives to improve system availability and MTTR.

Responsibilities include defining governance for observability, automating runbooks, and architecting telemetry for AWS/Azure with IaC.

Qualifications

  • Bachelor-level or higher in CS/IT or equivalent practical experience.
  • Experience in Infrastructure, Cloud, or SRE roles, ideally in regulated environments.
  • Hands-on with observability platforms (Datadog, Dynatrace, Splunk, ELK).
  • Automation with IaC and scripting (Terraform, Ansible, Python) and cloud observability tooling.

Responsibilities

  • Define enterprise observability architecture and governance across infra and apps.
  • Automate runbooks, self-healing workflows, and remediation using Python, Ansible, and Terraform.
  • Architect telemetry solutions for AWS and Azure and embed observability in CI/CD via IaC.
  • Collaborate with teams to drive resilience testing, chaos engineering, and incident advisories.

Skills

Infrastructure
Cloud
SRE
Regulated environments

Education

Qualification in Computer Science, Information Technology, or equivalent practical experience

Tools

Datadog
Dynatrace
Splunk
ELK Stack

Job description

Role Overview

The Observability Engineer leads the enterprise observability and service reliability strategy across infrastructure and application domains. This role drives proactive monitoring, automation, and resilience initiatives to enhance system availability, ensure regulatory compliance, and reduce mean time to resolution (MTTR).

Key Responsibilities
  • Observability Strategy & Governance: Define enterprise observability architecture aligned with operational resilience standards. Deploy and optimise full-stack observability platforms (metrics, logs, traces) and integrate them with ITSM and AIOps systems for predictive alerting.
  • Reliability Engineering & Automation: Implement SRE frameworks, define error budget policies, and codify operational reliability. Automate runbooks, self-healing workflows, and auto-remediation actions using Python, Ansible, and Terraform.
  • Cloud & Platform Observability: Architect and manage telemetry solutions for cloud-native workloads across AWS and Azure, embedding observability into landing zones and CI/CD pipelines via Infrastructure-as-Code (IaC).
  • Operational Excellence: Partner with cross-functional teams to conduct resilience testing, chaos engineering, and capacity validation. Maintain executive dashboards for operational risk indicators and act as a technical advisor during major incidents and audits.
Competencies and Qualifications
  • Technical Background: Qualification in Computer Science, Information Technology, or equivalent practical experience.
  • Domain Experience: Track record in Infrastructure, Cloud, or Site Reliability Engineering (SRE), with experience operating as an SRE subject matter expert, ideally within financial services or regulated environments.
  • Observability Platforms: Hands-on proficiency with tools such as Datadog, Dynatrace, Splunk, or ELK Stack.
  • Automation & Cloud: Expertise in Infrastructure-as-Code and scripting (Terraform, Ansible, Python) alongside cloud observability tooling (AWS CloudWatch/X-Ray, Azure Monitor/Log Analytics).
  • SRE & Compliance: Deep understanding of SRE principles (SLOs/SLAs, error budgets), financial sector operational resilience frameworks (e.g., MAS TRM, DORA), and automated remediation.
  • Certifications: Relevant certifications in observability platforms, IaC, cloud engineering (AWS/Azure), SRE, or ITIL are advantageous.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

Optimum Solutions (Singapore) Pte Ltd. • Singapore

On-site
SGD 90,000 - 130,000
Observability Engineer
Observability Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Datadog Admin
Datadog Admin

Optimum Solutions (Singapore) Pte Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

U3 INFOTECH PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Engineer
Engineer

OCBC Bank • Singapore

On-site
SGD 80,000 - 120,000
Associate VP (Support Engineering)
Associate VP (Support Engineering)

OCBC Bank • Singapore

On-site
SGD 80,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 180,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 250,000
Infrastructure Automation Engineer (AIOps & Agentic AI)
Infrastructure Automation Engineer (AIOps & Agentic AI)

Optimum Solutions (Singapore) Pte Ltd. • Singapore

On-site
SGD 90,000 - 150,000
Observability & Reliability Engineer
Observability & Reliability Engineer

Optimum Solutions (Singapore) Pte Ltd. • Singapore

On-site
SGD 90,000 - 130,000