Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
RecruitFirst Pte. Ltd invites applications for a Monitoring, Observability & Service Assurance role in Singapore.
You will manage centralized monitoring dashboards across applications, infrastructure, and cloud environments, and oversee 24/7 service health monitoring. Responsibilities include alert triage, cross‑team coordination, defining monitoring strategies, cost governance for cloud usage, and driving SRE practices with SLIs/SLOs.
Location: Tanjong Pagar
Working Hours: Office Hours
Salary: up $5,800 monthly basic
Manage and operate centralized monitoring dashboards and observability platforms across applications, databases, infrastructure (compute and storage), and network environments (on-premises and cloud) supporting 24/7 services.
Continuously track system health using metrics, logs, and alerts to proactively identify anomalies, performance degradation, and potential issues.
Respond to alerts and anomalies by conducting initial triage, impact assessment, and cross-system event correlation.
Coordinate and elevate issues to the appropriate teams (Application, Cloud/Infrastructure, Network, Database) to ensure timely resolution in line with SLAs.
Define and implement monitoring strategies, including frameworks, alert thresholds, escalation policies, and observability standards.
Collaborate with engineering teams to onboard systems into monitoring platforms and establish meaningful metrics and alerts.
Continuously review and refine monitoring frameworks to reduce noise and improve the signal-to-noise ratio.
Develop and maintain real-time service health dashboards for operational monitoring and reporting.
Track and analyze system, network, and cloud availability, performance trends, and recurring incident patterns.
Support reporting needs for service management, senior leadership, and key stakeholders.
Support major incident management by providing system visibility, diagnostics, and cross-team coordination.
Identify recurring issues and reliability gaps, driving improvements in system stability, monitoring coverage, and response times.
Monitor and manage cloud and infrastructure costs, including AWS usage (compute, storage, data transfer), as well as network and connectivity expenses.
Implement cost allocation, tagging strategies, and budget monitoring with alerting mechanisms.
Analyze cost drivers and identify optimization opportunities.
Prepare and present cost reports, dashboards, forecasts, and trend analyses.
Partner with engineering teams to optimize resource utilization, recommending rightsizing and cost-saving initiatives.
Ensure a balanced approach between cost efficiency, performance, and reliability.
Serve as the central coordination point across Application, Cloud/Infrastructure, Network, and Database teams to align monitoring insights with operational actions.
Drive initiatives to enhance observability, expand monitoring coverage, and automate alerting and response workflows.
Contribute to the adoption of Site Reliability Engineering (SRE) practices, including SLIs, SLOs, and error budgets.
Maintain documentation for monitoring architecture and dashboards, alerting rules and escalation procedures, cost governance models, and reports.
Participate in major incident response and critical service monitoring.
Provide after-hours support, including weekends and public holidays, as required.
Training in Computer Science, Information Technology, Engineering, or a related field.
3 to 5 years of experience in IT operations, system monitoring, NOC, service assurance, or cloud/infrastructure operations.
Hands-on experience with monitoring and observability platforms.
Experience working in hybrid environments (on-premises and AWS cloud).
Strong understanding of system and network monitoring concepts, as well as application and infrastructure health metrics.
Proficiency with tools such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or similar platforms.
Ability to analyze and interpret logs, metrics, and alerts effectively.
Experience with AWS Cost Explorer, budgeting, and tagging strategies, with a good understanding of cloud cost structures and optimization techniques.
Familiarity with ITIL processes (Incident, Problem, and Change Management), service level management, and observability practices.
AWS certifications (Associate level or above) or AWS FinOps Certified Practitioner are preferred.
Familiarity with ITIL processes (Incident, Problem, Change Management), Service Level Management and Observability principles
AWS Certification (Associate level or above) or AWS FinOps Certified Practitioner preferred
Strong analytical thinking and problem-solving abilities.