An application made for this job — a tailored resume and cover letter that speak straight to the posting.
SEARCH INDEX PTE. LTD. in Singapore seeks an IT Monitoring & Service Assurance specialist to maintain monitoring platforms across applications, databases, infrastructure, and cloud environments.
You will monitor health via logs and metrics, develop dashboards, support AWS cloud costs, and help drive reliability with SRE practices. The role requires 3–5 years in IT operations and strong analytical skills.
Monitoring & Service Assurance
Cloud FinOps & Governance
Incident & Service Operation
Manage and maintain monitoring and observability platforms across applications, databases, infrastructure, and network environments (on-premise and cloud).
Monitor system health through logs, metrics, and alerts to identify issues, perform incident triage, and coordinate timely resolution with relevant team.
Develop and maintain dashboards and reports to monitor service availability, performance trends, and operational insights.
Support cloud cost monitoring and governance initiatives, including cost tracking, tagging strategies, and optimization opportunities.
Drive continuous improvements in monitoring coverage, automation, and operational processes to enhance service reliability and efficiency.
Participate in incident management, maintain documentation, and provide after-hours support when required.
Degree/Diploma in Computer Science, Information Technology, Engineering, or related discipline.
Minimum 3–5 years of experiences in IT operations, infrastructure support, service assurance, NOC, or cloud environments.
Hands-on experience with monitoring and observability platforms such as CloudWatch, Grafana, Prometheus, Splunk, ELK Stack, or similar tools.
Experience working in hybrid environments (on-premise and AWS cloud).
Strong understanding of infrastructure, system, network, and application monitoring concepts.
Familiarity with AWS Cost Explorer, cloud budgeting, tagging strategies, and cost optimization practices.
Knowledge of ITIL processes including Incident, Problem, and Change Management.
Exposure to SRE practices and service reliability principles is advantageous.
AWS Associate Certification and/or AWS FinOps Certified Practitioner preferred.
Strong analytical and troubleshooting skills with the ability to correlate events across complex systems.
Strong communication and stakeholder management abilities.
Self-driven, proactive, and able to work independently in a fast-paced environment.
Willing to provide after-hours support when required.