IT Infrastructure Engineer

GMP RECRUITMENT SERVICES (S) PTE LTD

Singapore

On-site

SGD 60,000 - 110,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

GMP RECRUITMENT SERVICES (S) PTE LTD in Singapore seeks an IT operations professional to manage centralized monitoring and observability across applications, databases, infrastructure, networks, and cloud to ensure 24/7 service availability.

You will monitor health with metrics, logs, and alerts; perform incident coordination, design monitoring strategies, support cost optimization, and drive continuous improvement through automation and documentation, including after-hours deployments.

Qualifications

  • 3–5 years of experience in IT operations, NOC, service assurance, system monitoring, or cloud/infrastructure operations.
  • Hands-on experience with monitoring and observability tools such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or equivalent platforms.
  • Strong understanding of hybrid infrastructure (on-premises and AWS), system/network monitoring, application performance, and log/metric analysis.
  • Experience with AWS cost management, including Cost Explorer, budgeting, tagging strategies, and cloud cost optimization practices.
  • Familiarity with ITIL processes (Incident, Problem, and Change Management); AWS Associate-level certification or AWS FinOps Certified Practitioner is preferred.
  • Willingness to support after-hours operations, including deployments, maintenance, and incident response.

Responsibilities

  • Manage and operate centralized monitoring and observability platforms across applications, databases, infrastructure, networks, and cloud environments to ensure 24/7 service availability.
  • Monitor system health using metrics, logs, and alerts; proactively identify anomalies, performance issues, and service degradation.
  • Perform alert triage, impact assessment, and incident coordination, escalating issues to the appropriate technical teams to meet SLA requirements.
  • Design and enhance monitoring strategies, dashboards, alerting frameworks, and observability standards to improve service visibility and reduce alert noise.
  • Support major incident management by providing diagnostics, cross-team coordination, and driving service reliability improvements through trend analysis and root cause identification.
  • Monitor and optimize cloud and infrastructure costs, implementing tagging, budgeting, cost allocation, and identifying opportunities for cost savings.
  • Develop operational and cost reports, dashboards, and forecasts to support service management, leadership, and operational decision-making.
  • Drive continuous improvement by expanding monitoring coverage, automating observability processes, maintaining documentation, and supporting after-hours operational activities.

Skills

IT operations
NOC / service assurance
Incident coordination

Tools

CloudWatch
Grafana
Prometheus
Splunk
ELK Stack

Job description

Responsibilities
  • Manage and operate centralized monitoring and observability platforms across applications, databases, infrastructure, networks, and cloud environments to ensure 24/7 service availability.

  • Monitor system health using metrics, logs, and alerts; proactively identify anomalies, performance issues, and service degradation.

  • Perform alert triage, impact assessment, and incident coordination, escalating issues to the appropriate technical teams to meet SLA requirements.

  • Design and enhance monitoring strategies, dashboards, alerting frameworks, and observability standards to improve service visibility and reduce alert noise.

  • Support major incident management by providing diagnostics, cross-team coordination, and driving service reliability improvements through trend analysis and root cause identification.

  • Monitor and optimize cloud and infrastructure costs, implementing tagging, budgeting, cost allocation, and identifying opportunities for cost savings.

  • Develop operational and cost reports, dashboards, and forecasts to support service management, leadership, and operational decision-making.

  • Drive continuous improvement by expanding monitoring coverage, automating observability processes, maintaining documentation, and supporting after-hours operational activities.

Requirements
  • 3–5 years of experience in IT operations, NOC, service assurance, system monitoring, or cloud/infrastructure operations.

  • Hands-on experience with monitoring and observability tools such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or equivalent platforms.

  • Strong understanding of hybrid infrastructure (on-premises and AWS), system/network monitoring, application performance, and log/metric analysis.

  • Experience with AWS cost management, including Cost Explorer, budgeting, tagging strategies, and cloud cost optimization practices.

  • Familiarity with ITIL processes (Incident, Problem, and Change Management); AWS Associate-level certification or AWS FinOps Certified Practitioner is preferred.

    • Willingness to support after-hours operations, including deployments, maintenance, and incident response.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Engineer, IT Infrastructure (up $5,800)
Senior Engineer, IT Infrastructure (up $5,800)

RecruitFirst Pte. Ltd • Singapore

On-site
SGD 59,000 - 71,000
IT Infrastructure Engineer (Monitoring & Operations)
IT Infrastructure Engineer (Monitoring & Operations)

SEARCH INDEX PTE. LTD. • Singapore

On-site
SGD 67,000 - 100,000
Observability / ITSM Engineer
Observability / ITSM Engineer

ENGGSOL PTE. LTD. • Singapore

On-site
SGD 60,000 - 90,000
IT Infrastructure Operations Manager
IT Infrastructure Operations Manager

ariston services pte. ltd. • Singapore

On-site
SGD 120,000 - 160,000
IT Infrastructure Engineer
IT Infrastructure Engineer

ITCAN PTE. LIMITED • Singapore

On-site
SGD 90,000 - 140,000
Technical Operations Engineer
Technical Operations Engineer

International SOS • Singapore

On-site
SGD 120,000 - 190,000
Observability Engineer (Logging & Monitoring)
Observability Engineer (Logging & Monitoring)

Accenture Southeast Asia • Singapore

On-site
SGD 90,000 - 140,000
Operations Support Engineer (ITSM, AWS) - #1563
Operations Support Engineer (ITSM, AWS) - #1563

JOBSTER PRIVATE LTD. • Singapore

On-site
SGD 60,000 - 90,000
IT Infrastructure Executive – Infrastructure Operations
IT Infrastructure Executive – Infrastructure Operations

SBS Transit Limited • Singapore

Hybrid
SGD 90,000 - 130,000
Technical Operations Engineer
Technical Operations Engineer

International SOS group • Singapore

Hybrid
SGD 90,000 - 150,000