L1 - Azure Operations & Monitoring Engineer
L1 – Azure Operations & Monitoring Engineer
Scope Summary
The L1 Operations & Monitoring Engineer will be responsible for proactive monitoring, operational governance, first-level troubleshooting, incident management, alert handling, deployment validation support, and day-to-day infrastructure operational stability across Azure and integrated cloud platforms. The role focuses on ensuring infrastructure availability, operational continuity, alert governance, SLA adherence, and escalation management across production and non-production environments.
Key Scope of Work
Infrastructure & Platform Monitoring
- 9x6 monitoring of Azure VMs, AKS clusters, networking components, ADX, HDInsight, storage services, CDN services, and cloud platform availability.
- Monitor infrastructure health, CPU, memory, disk utilization, network latency, storage performance, SSL validity, and service availability.
- Monitor AKS pods, nodes, ingress controllers, deployments, namespaces, and cluster health.
- Monitor database platform health for ScyllaDB, MongoDB, Redis, and Azure SQL environments.
- Monitor Fastly CDN, JioCDN, Firebase, GCP services, and integrated cloud services.
- Handle first-level troubleshooting, service restarts, pod restarts, alert acknowledgement, and issue tracking.
- Monitor CloudXP, Azure Monitor, Log Analytics, and centralized observability alerts.
- Coordinate and escalate critical alerts and incidents to L2/L3 teams.
- Track incidents until closure and maintain operational governance.
Backup, DR & Operational Governance
- Monitor backup jobs, DR status, recovery readiness, and infrastructure health.
- Validate deployment completion and support release validation activities.
- Prepare daily operational reports, health reports, incident summaries, and availability dashboards.
- Ensure ticket management, SLA adherence, and operational response tracking.
Security & Compliance Monitoring
- Monitor WAF alerts, Azure Defender alerts, failed login attempts, SSL expiry alerts, and infrastructure security events.
- Support security monitoring and operational governance across production platforms.
On-Call & Operational Support
- Participate in operational support activities and escalation coordination during critical incidents.
- Support after-hours incident escalation coordination as per defined operational governance.
Priority Scope Item
- P1 9x6 Infrastructure & Platform Monitoring
- P1 AKS Pod, Node & Deployment Monitoring
- P1 Backup & DR Monitoring
- P1 Security Monitoring (WAF, Defender, SSL, Failed Logins)
- P1 Ticket Management & SLA Tracking
- P1 Daily Operational Reporting & Availability Reporting
- P1 On-Call & After-Hours Incident Coordination