Job Overview
You will be deployed to our client who is a leading government technology agency driving digital transformation and innovation in the healthcare sector. This role will focus on managing and scaling an Elastic SaaS-based observability platform to enhance infrastructure resiliency. The system engineer designs, deploys, and optimizes telemetry pipelines for logs, metrics, and traces using tools like Elastic Stack, Logstash, and supporting infrastructure. Responsibilities emphasize platform availability (99.99% SLO), security, performance, and self-service capabilities for operations teams.
Responsibilities
- Responsible for maintaining the team’s platform, implementing effective monitoring, metrics, logging, and visualisation systems.
- Work with vendors to design and develop the platform comprising of monitoring, metrics, logging systems and AI/ML functionalities.
- Collaborate with internal teams to onboard existing infrastructure devices to observability platform.
- Conceptualize and implement early anomaly detection (reduction of mean-time issue identification), pattern analysis, self-healing, infrastructure resizing, noise reduction and outage prediction.
- Develop and maintain visualization in observability tools, providing single pane views for end user experience for infrastructure and security.
- Provide expertise support to the end users for report customisation and dashboarding
Requirements
- Bachelor in Computer Engineering, Computer Science, or equivalent
- Min 3 years of experience in Enterprise-level infrastructure environment and 1 year of experience in Enterprise-level monitoring and/or Logging Technologies.
- Hands-on experience with ELK (Elasticsearch) and Red Hat Enterprise Linux (RHEL).
- Hands-on experience with modern observability tools and frameworks is preferred but open to candidates with prior experience in monitoring solutions who are looking to transition into modern observability practices
- Strong foundation in monitoring and troubleshooting
- Have good knowledge in infrastructure monitoring and logging technologies.
- Have working knowledge in working with programming and scripting languages.
- Have working knowledge in infrastructure automation.
- Proficiency in PHEL/Linux/Unix, scripting (Ruby, Shell, Painless)
- Strong problem-solving skills and a proactive attitude towards learning new technologies
- Excellent communication, documentation and teamwork skills, with the ability to work effectively in a collaborative environment
- Able to support a 24/7 operational environment and respond to support requirements when needed.
Only shortlisted candidates will be responded to, therefore if you do not receive a reply within 14 days please accept this as notification that you have not been shortlisted.