Job Summary
We are looking for an experienced Azure L2 Support Engineer with strong expertise in Azure cloud operations, ITIL-based support, incident/problem management, and data pipeline monitoring. The candidate will be responsible for providing L2 production support, troubleshooting Azure services and data pipelines, resolving incidents within SLA, and coordinating with L3/Development teams for complex issues.
Mandatory Skills
- 613 years of IT experience with relevant experience in Azure Production Support.
- Strong hands-on experience in L2 Application/Cloud Support.
- Strong understanding of ITIL processes.
- Experience in Incident Management, Problem Management, Change Management and Service Request Management.
- Experience monitoring and troubleshooting Azure Data Pipelines.
- Strong knowledge of Azure Data Factory (ADF).
- Good knowledge of SQL for data validation and troubleshooting.
- Experience in analyzing application/data pipeline failures and logs.
- Strong understanding of SLA, priority, severity and escalation processes.
- Experience working in 24x7 production support / on-call environments is preferred.
Key Responsibilities
- Provide L2 production support for Azure-based applications and data platforms.
- Monitor Azure services and data pipelines and identify failures proactively.
- Troubleshoot failed or delayed Azure Data Factory pipelines.
- Analyze pipeline logs, error messages, dependencies, triggers, activities, and data movement issues.
- Perform initial root-cause analysis and resolve incidents within defined SLA.
- Validate data and troubleshoot issues using SQL queries.
- Monitor scheduled jobs, batch processes, data loads, and downstream dependencies.
- Restart/reprocess failed pipelines based on approved operational procedures.
- Identify recurring issues and contribute to Problem Management / RCA activities.
- Create and maintain incident, problem, change, and knowledge-management documentation.
- Follow ITIL processes and operational procedures.
- Coordinate with L3 developers, data engineers, architects, and infrastructure teams for complex issues.
- Participate in change implementation, validation, and post-deployment checks.
- Perform production health checks and provide operational status updates.
- Maintain support dashboards, incident reports, and SLA metrics.
- Participate in on-call/shift support as required.