Dear Candidate,
We are hiring for Observability Engineer role for MUmbai Location
If any profile is matching please do share me your profile.
Overview:
The Operational Resilience drive gives an opportunity to improve upon current infrastructure safeguards. With focus on protecting our most important assets, this role will drive a business down view of service availability giving everyone a true state of the current service levels and provide comfort that we are extending Operational Resilience principles into our extensive monitoring suite.
Role / Principal Accountabilities: -
- Build the plan based on the in-house Observability stack.
- Deliver a plan to meet Operational Resilience monitoring requirements based on IBS/ CIF business functions: Settlement/Collateral etc,
- Produce a business dashboard to give visibility of business-critical services and show real-time service status.
- Align with the in-house Observability team to ensure best practices are achieved.
- Align with ServiceNow Configuration Management to ensure any missing assets are added to ServiceNow.
- Work with Operational Resilience and infrastructure management to gain buy-in and visibility of the project.
- Lead a team of monitoring experts.
- Use SRE monitoring principles.
- Align with SRE management.
- Align with Infra management.
- Contribute to Observability standards and procedures.
- Work with the SRE team to optimize the operation model Secondary.
- Align with Application teams to ensure their systems are integrated into the overall service dashboard.
Skills & Experience Required: -
- Strong communication – presenting and documentation.
- Team management and stakeholder skills.
- Experience of driving change in a large organization.
- Demonstrate understanding of large enterprise systems and the technologies on which they are built.
- Strong analytical skills and a solid understanding from a Production Support monitoring point of view .
- In-depth understanding of the technical aspects of delivering monitoring integration.
Knowledge Required: -
- Background in Observability Engineering covering Metrics, Logs and Traces and the interaction between telemetry.
- Strong Experience with Observability tools such as Open Telemetry, Grafana UI, Mimir, Loki, Tempo, Grafana Agent, Prometheus.
- Appreciation of SLI/SLOs for measuring and reporting on service levels.
- Awareness of notification frameworks for escalation automation.
- Experience of DevOps tools i.e. Jenkins, Ansible Tower, GIT and build & deploy pipelines.
- Exposure to Kubernetes, containers & micro-service concepts (not essential, but beneficial)