Position Specific Requirements
- Incident handling, Issue analysis, Investigate & Gather necessary data/insights
- Incident escalation based on workflows
- Own L2 incident handling: triage, isolate, restore, and document; run bridges for P1/P2 and steer the right resolver group.
- Monitor & act on CloudWatch/Datadog/ELK dashboards & auto alerts using runbooks; tune noise vs. signal.
- Deep dive production issues across Lambda/API Gateway/OpenSearch/Athena; create actionable debug artifacts (timelines, queries, log snippets, correlation IDs).
- Partner with L3: supply high quality inputs (repro steps, traces, request samples, dashboards), verify fixes, and drive post incident actions.
- Operational hygiene: maintain runbooks, workflows, SOPs, and knowledge articles; contribute to SLA reporting and weekly ops reviews.
- Cross geo coordination: collaborate with business, SupportDesk, QA, SecOps, vendors; close tickets end to end.
- Continuous improvement: raise kaizens to reduce toil (alerts, auto remediations, dashboards, cost/perf optimizations in Athena/OpenSearch).
- In case of priority incidents , bridge the calls with multiple stakeholders and ensure resolution stakeholder is identified and assigned to incident
- Monitoring Dashboard of cloud applications
- Monitoring and acting based on the auto-alerts generated from system using runbooks
- Monitoring and Handling of the slack Datadog alerts
- Report creation and publishing to required stake holders
- Responsible for maintaining the procedures/run books, workflows, work instructions used by the ops team
- Fulfil any ad-hoc data or report request queries from different functional groups
Shifts
- Rotational 247 (multiple shifts), including weekends oncall / holidays support as rostered.
Position Specific Skills
- 3+ years of experience in Level 2 Application support
- Expertise in application and backend issue triage & analysis. Should have analyzed failures, error logs, done basic analysis, gather inputs and coordinate with Level 3
- Own day-to-day product operations for cloud-hosted applications, including proactive monitoring, incident triaging, outage management, and resolution while ensuring adherence to defined SLAs.
- Demonstrate strong analytical capabilities to diagnose issues, drive incidents to closure, and adapt to evolving tools and techniques during real-time incident response.
- Having worked upon support layers L1, L2, L2.5, worked closely with the engineering L3 teams, Product teams and the Business Stakeholders
- Troubleshooting knowledge in AWS, containerized workloads, and distributed microservices architecture.
- Understanding the root causes, restoration of the services and final reporting
- Monitor application and infrastructure using Datadog/Logicmonitor/Newrelic/Grafana or any monitoring tool
- Usage of multiple tools and techniques to analyse the production issues across Applications, backend, data and infrastructure layers
- Correlate logs, traces and metrics to isolate the failures across multiple services, understand the impact on business and generate high quality artifacts aiding L3 investigations
- Support and lead release management, operational readiness, deployment support and postproduction deployment validations
- Collaborate with the Stakeholders, Project teams and Vendors, handle the incident triages and meetings ensuring SLA adherences, escalation flows and operational excellence
- Good command in Microsoft Office suite
- Willing to work in rotational based 24X7 multiple Shifts
- Willing to learn and implement automation or AIOps Features
Tools (Related should be fine too)
Monitoring and Observability
- Datadog, ElasticSearch ELK / Open Search Dashboards, AWS Cloudwatch
Cloud & Infra
- AWS, Lambda, API Gateway, Open Search, Athena, Web services, API Management, AWS Cognito, EKS, SQS
Incident Management ITSM tools
Databases
- PostgreSQL, Mongo No SQL, SQL