Senior Engineer, IT Infrastructure (up $5,800)

RecruitFirst Pte. Ltd

Singapore

On-site

SGD 59,000 - 71,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

RecruitFirst Pte. Ltd invites applications for a Monitoring, Observability & Service Assurance role in Singapore.

You will manage centralized monitoring dashboards across applications, infrastructure, and cloud environments, and oversee 24/7 service health monitoring. Responsibilities include alert triage, cross‑team coordination, defining monitoring strategies, cost governance for cloud usage, and driving SRE practices with SLIs/SLOs.

Qualifications

  • 3–5 years of experience in IT operations, system monitoring, NOC or cloud/infrastructure operations.
  • Hands-on experience with monitoring and observability platforms.
  • Experience in hybrid on‑premises and AWS cloud environments.
  • Strong understanding of monitoring concepts and reliability metrics.
  • Proficient with leading monitoring tools (CloudWatch, Grafana, Prometheus, Splunk, ELK).
  • Knowledge of ITIL practices and cost governance in cloud environments.

Responsibilities

  • Manage centralized monitoring dashboards across applications, databases, infrastructure, and networks (on‑premises and cloud).
  • Continuously track system health using metrics, logs and alerts to detect anomalies.
  • Respond to alerts with initial triage and impact assessment; coordinate cross‑team resolution.
  • Define monitoring strategies, thresholds, escalation policies and observability standards.
  • Develop real‑time service health dashboards and report on availability and performance trends.
  • Support major incident management with diagnostics and cross‑team coordination.
  • Monitor and manage cloud and infrastructure costs; implement tagging and budgeting alerts.
  • Produce cost reports and collaborate with engineering to optimize resource use.
  • Act as central coordinator across Application, Cloud/Infrastructure, Network, and Database teams.
  • Drive SRE evolution with SLIs, SLOs and error budgets; expand monitoring coverage.

Skills

IT operations
System monitoring
NOC
Service assurance
Cloud/infrastructure operations

Education

Bachelor's degree in Computer Science/IT/Engineering

Tools

CloudWatch
Grafana
Prometheus
Splunk
ELK Stack

Job description

Location: Tanjong Pagar
Working Hours: Office Hours
Salary: up $5,800 monthly basic

Monitoring, Observability & Service Assurance
  • Manage and operate centralized monitoring dashboards and observability platforms across applications, databases, infrastructure (compute and storage), and network environments (on-premises and cloud) supporting 24/7 services.

  • Continuously track system health using metrics, logs, and alerts to proactively identify anomalies, performance degradation, and potential issues.

Alerting, Triage & Coordination
  • Respond to alerts and anomalies by conducting initial triage, impact assessment, and cross-system event correlation.

  • Coordinate and elevate issues to the appropriate teams (Application, Cloud/Infrastructure, Network, Database) to ensure timely resolution in line with SLAs.

Monitoring Strategy, Design & Governance
  • Define and implement monitoring strategies, including frameworks, alert thresholds, escalation policies, and observability standards.

  • Collaborate with engineering teams to onboard systems into monitoring platforms and establish meaningful metrics and alerts.

  • Continuously review and refine monitoring frameworks to reduce noise and improve the signal-to-noise ratio.

Service Health & Reporting
  • Develop and maintain real-time service health dashboards for operational monitoring and reporting.

  • Track and analyze system, network, and cloud availability, performance trends, and recurring incident patterns.

  • Support reporting needs for service management, senior leadership, and key stakeholders.

Incident Support & Service Reliability
  • Support major incident management by providing system visibility, diagnostics, and cross-team coordination.

  • Identify recurring issues and reliability gaps, driving improvements in system stability, monitoring coverage, and response times.

FinOps, Cost Monitoring & Governance
  • Monitor and manage cloud and infrastructure costs, including AWS usage (compute, storage, data transfer), as well as network and connectivity expenses.

  • Implement cost allocation, tagging strategies, and budget monitoring with alerting mechanisms.

  • Analyze cost drivers and identify optimization opportunities.

Cost Reporting & Optimisation
  • Prepare and present cost reports, dashboards, forecasts, and trend analyses.

  • Partner with engineering teams to optimize resource utilization, recommending rightsizing and cost-saving initiatives.

  • Ensure a balanced approach between cost efficiency, performance, and reliability.

Cross-Team Coordination
  • Serve as the central coordination point across Application, Cloud/Infrastructure, Network, and Database teams to align monitoring insights with operational actions.

Continuous Improvement & SRE Evolution
  • Drive initiatives to enhance observability, expand monitoring coverage, and automate alerting and response workflows.

  • Contribute to the adoption of Site Reliability Engineering (SRE) practices, including SLIs, SLOs, and error budgets.

Documentation
  • Maintain documentation for monitoring architecture and dashboards, alerting rules and escalation procedures, cost governance models, and reports.

Operational Support
  • Participate in major incident response and critical service monitoring.

  • Provide after-hours support, including weekends and public holidays, as required.

Requirements:
  • Training in Computer Science, Information Technology, Engineering, or a related field.

  • 3 to 5 years of experience in IT operations, system monitoring, NOC, service assurance, or cloud/infrastructure operations.

  • Hands-on experience with monitoring and observability platforms.

  • Experience working in hybrid environments (on-premises and AWS cloud).

  • Strong understanding of system and network monitoring concepts, as well as application and infrastructure health metrics.

  • Proficiency with tools such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or similar platforms.

  • Ability to analyze and interpret logs, metrics, and alerts effectively.

  • Experience with AWS Cost Explorer, budgeting, and tagging strategies, with a good understanding of cloud cost structures and optimization techniques.

  • Familiarity with ITIL processes (Incident, Problem, and Change Management), service level management, and observability practices.

  • AWS certifications (Associate level or above) or AWS FinOps Certified Practitioner are preferred.

  • Familiarity with ITIL processes (Incident, Problem, Change Management), Service Level Management and Observability principles

  • AWS Certification (Associate level or above) or AWS FinOps Certified Practitioner preferred

  • Strong analytical thinking and problem-solving abilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Operations Engineer (Cloud & Monitoring)
Infrastructure Operations Engineer (Cloud & Monitoring)

Search Index Pte Ltd • Singapore

On-site
SGD 70,000 - 110,000
Cloud Engineer (Monitoring / AWS)
Cloud Engineer (Monitoring / AWS)

SEARCH INDEX PTE. LTD. • Singapore

On-site
SGD 70,000 - 110,000
Infrastructure Operations Engineer (Cloud & Monitoring)
Infrastructure Operations Engineer (Cloud & Monitoring)

SEARCH INDEX PTE. LTD. • Singapore

On-site
SGD 60,000 - 90,000
IT Infrastructure Engineer (Monitoring & Operations)
IT Infrastructure Engineer (Monitoring & Operations)

SEARCH INDEX PTE. LTD. • Singapore

Hybrid
SGD 67,000 - 100,000
IT Infrastructure Engineer
IT Infrastructure Engineer

GMP RECRUITMENT SERVICES (S) PTE LTD • Singapore

On-site
SGD 60,000 - 110,000
Cloud Operations Engineer (Monitoring / AWS)
Cloud Operations Engineer (Monitoring / AWS)

SEARCH INDEX PTE. LTD. • Singapore

On-site
SGD 60,000 - 80,000
Cloud Operations Engineer (Monitoring)
Cloud Operations Engineer (Monitoring)

SEARCH INDEX PTE. LTD. • Singapore

Hybrid
SGD 80,000 - 110,000
IT Infrastructure Engineer (Monitoring & Operations)
IT Infrastructure Engineer (Monitoring & Operations)

SEARCH INDEX PTE LTD • Singapore

On-site
SGD 60,000 - 90,000
Cloud Engineer
Cloud Engineer

WORLD PARTNERS SOLUTION PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Observability Engineer
Observability Engineer

GOLDTECH RESOURCES PTE LTD • Singapore

On-site
SGD 120,000 - 180,000