Observability Engineer

GOLDTECH RESOURCES PTE LTD

Singapore

On-site

SGD 120,000 - 180,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GOLDTECH RESOURCES PTE LTD is seeking a Senior Observability Engineer to strengthen monitoring across hybrid infrastructure and cloud environments. You will improve system visibility, support platform availability and drive automation across IT operations.

You will work with cross-functional teams to define metrics, reliability indicators and operational reporting, and you will implement monitoring solutions using SolarWinds, Datadog, Azure Monitor, Grafana or Splunk, while contributing to

Qualifications

  • 8+ years in systems administration, SysOps, or observability.
  • Experience in enterprise IT environments.
  • Hands-on with Windows Server and Linux.
  • Hands-on with monitoring platforms (SolarWinds, Datadog, Azure Monitor, Grafana, Splunk).
  • Experience with Azure IaaS/PaaS and cloud security.
  • Exposure to containers and Kubernetes.
  • Automation scripting using PowerShell, Bash, or Python.
  • Experience with IaC tools (Ansible, Terraform).
  • Knowledge of ITIL and incident management.

Responsibilities

  • Manage and support monitoring across on-prem and cloud.
  • Configure monitoring for infrastructure, apps, logs, events, service health.
  • Develop dashboards and operational views for performance and reliability.
  • Establish health checks, capacity trends, baselines.
  • Tune alerts and reduce unnecessary alerts.
  • Automate provisioning, health checks, reporting and remediation.
  • Contribute to incident response, RCA, and post-incident reviews.
  • Support security, change management and compliance.

Skills

Observability
SysOps
SRE
Incident management
Automation scripting
DevOps concepts
Stakeholder comms

Tools

SolarWinds
Datadog
Azure Monitor
Grafana
Splunk
PowerShell
Bash
Python
Ansible
Terraform

Job description

Role Overview

We are looking for an experienced Senior Observability Engineer to support and enhance enterprise monitoring and observability capabilities across hybrid infrastructure and cloud environments.

The role will focus on maintaining platform availability, improving system visibility, strengthening incident response and introducing greater automation across IT operations. You will work across infrastructure, applications and cloud services to improve monitoring coverage, alert effectiveness, root cause analysis and overall service reliability.

Key Responsibilities
Observability & Monitoring
  • Manage and support monitoring and observability platforms across on-premises and cloud environments.
  • Configure and maintain monitoring solutions for infrastructure, applications, logs, events and service health.
  • Develop dashboards and operational views covering application performance, infrastructure health, service trends and incident impact.
  • Establish health checks, availability monitoring, capacity trends and operational baselines.
  • Improve monitoring effectiveness through alert tuning, event correlation and reduction of unnecessary alerts.
  • Implement monitoring solutions using tools such as SolarWinds, Datadog, Azure Monitor, Grafana, Splunk or similar platforms.
  • Work with technical teams to define service metrics, reliability indicators and operational reporting.
Reliability & Incident Management
  • Support the investigation and resolution of incidents affecting critical infrastructure, applications and IT services.
  • Conduct root cause analysis and recommend corrective and preventive actions.
  • Participate in post-incident reviews and identify opportunities to improve system reliability.
  • Maintain escalation procedures, recovery processes and operational runbooks.
  • Support high availability and disaster recovery planning, testing and continuous improvement.
Automation & Operational Improvement
  • Automate operational activities such as provisioning, configuration, health checks, reporting and remediation.
  • Develop and maintain automation using PowerShell, Bash, Python, Ansible, Terraform or equivalent technologies.
  • Support Infrastructure as Code and configuration management practices.
  • Maintain reusable scripts, templates, workflows and automation runbooks using appropriate version control.
  • Explore AI-assisted operational capabilities for log analysis, alert correlation, incident triage and operational reporting.
  • Review and validate AI-assisted recommendations or scripts before implementation.
Security, Compliance & Change Management
  • Apply security and configuration standards across supported infrastructure and cloud environments.
  • Support vulnerability remediation, system hardening, patching, access controls and secure configuration.
  • Ensure monitoring and infrastructure changes follow established change and release management processes.
  • Maintain appropriate operational records to support security reviews and audits.
  • Support environments operating under standards such as PCI-DSS and ISO/IEC 27001 .
IT Service Management & Documentation
  • Follow ITIL-based incident, problem, change, request and configuration management processes.
  • Maintain accurate system documentation, operational procedures, monitoring configurations and service dependencies.
  • Maintain configuration and service information within CMDB or approved inventory systems.
  • Develop and maintain operational runbooks and technical procedures .
  • Work closely with infrastructure, application, cybersecurity, service management and vendor teams.
  • Promote operational standards and knowledge sharing across Windows, Linux, Azure and container environments.
Requirements
  • 8+ years of experience in Systems Administration, SysOps, Infrastructure Operations, Observability, Site Reliability Engineering or a related area.
  • Strong experience supporting enterprise or business-critical IT environments.
  • Hands-on knowledge of Windows Server and Linux environments.
  • Experience with enterprise monitoring and observability platforms such as SolarWinds, Datadog, Azure Monitor, Grafana or Splunk.
  • Good understanding of Microsoft Azure , including IaaS, PaaS and cloud security.
  • Exposure to containers and Kubernetes .
  • Hands-on scripting or automation experience using PowerShell, Bash or Python .
  • Experience with automation and Infrastructure as Code tools such as Ansible or Terraform would be advantageous.
  • Understanding of CI/CD, infrastructure automation and modern DevOps practices.
  • Familiarity with ITIL processes and enterprise incident management.
  • Knowledge of security and compliance frameworks such as PCI-DSS and ISO/IEC 27001 .
  • Strong troubleshooting and root cause analysis capabilities.
  • Able to communicate effectively with both technical and non-technical stakeholders.
  • Comfortable supporting critical systems and participating in weekend, evening or on-call coverage when required.
  • Preferred Certif
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Engineer
Observability Engineer

U3 SOLUTIONS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Observability Engineer (Logging & Monitoring)
Observability Engineer (Logging & Monitoring)

Accenture Southeast Asia • Singapore

On-site
SGD 90,000 - 140,000
IT Infrastructure Engineer
IT Infrastructure Engineer

GMP RECRUITMENT SERVICES (S) PTE LTD • Singapore

On-site
SGD 60,000 - 110,000
M01 - Observability Engineer
M01 - Observability Engineer

FPT ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
M01 - Observability Engineer
M01 - Observability Engineer

FPT Asia Pacific • Singapore

On-site
SGD 90,000 - 120,000
Observability Engineer - SRE / Cloud Data
Observability Engineer - SRE / Cloud Data

NEXBRIDGE RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
Associate Systems Engineer (Observability) - Central Infra
Associate Systems Engineer (Observability) - Central Infra

Synapxe • Singapore

On-site
SGD 60,000 - 90,000
Systems Engineer (ELK Observability) - Infra Operations
Systems Engineer (ELK Observability) - Infra Operations

SYNAPXE PTE. LTD. • Singapore

On-site
SGD 90,000 - 150,000
IT Infrastructure Engineer
IT Infrastructure Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 140,000
LiveOps Cloud Engineer #2209
LiveOps Cloud Engineer #2209

EduCare Specialist Services • Singapore

On-site
SGD 90,000 - 150,000