ZYBISYS CONSULTING SERVICES LLP | Full time
L1 Operations Engineer
Bangalore jayanagar, India | Posted on 08/04/2026
EmploymentType: Full-Time|Shift-Based (24x7 Rotational)
About Zybisys
Zybisys is a technology companythat helps banks, financial institutions, and FinTech businesses build and runsecure, reliable, and high-performance technology platforms. We work closelywith some of India's leading stock brokers to manage their cloudinfrastructure, cybersecurity, platform operations, and observability. Withdeep expertise in the Capital Markets domain, we focus on simplifying complextechnology, improving operational resilience, and helping our customers innovatewith confidence.
Job Description
The L1 Operations Engineer isthe first line of defence in Zybisys's 24x7 managed cloud operations. This roleis responsible for continuous monitoring of multi-region Azure cloud and hybridon-premises infrastructure, timely triage of alerts, first-line incidentresponse, and ensuring all issues are accurately logged, prioritised, andescalated within SLA thresholds.
This is a shift-basedoperational role requiring strong attention to detail, disciplined runbookexecution, and clear communication during incidents. The ideal candidate istechnically curious, process-oriented, and comfortable working across cloudmonitoring tools, ITSM platforms, and infrastructure dashboards in a fast-pacedmanaged services environment.
Key Responsibilities
Infrastructure Monitoring &Alert Management
- Perform continuous 24x7monitoring of Azure cloud and hybrid datacenter infrastructure usingPrometheus, Grafana, Azure Monitor, and related observability dashboards.
- Triage incoming alerts — assessseverity, validate against known patterns, and determine whether to resolve atL1 or elevate to L2 within defined SLA thresholds.
- Execute approved runbooks andSOPs for all known alert categories; document actions taken for every incidentwith accurate timestamps and observations.
- Monitor health and availabilityof compute (VMs), storage, network links, VPN tunnels, and platform servicesacross cloud and on-premises environments.
- Track sFlow and NetFlowdashboards for network traffic anomalies; flag unusual patterns to the L2 teamfor deeper investigation.
- Log all incidents, servicerequests, and alerts in the ITSM platform with complete and accurate details —symptoms, affected components, priority, and initial actions taken.
- Update ticket status throughoutthe incident lifecycle; ensure no incident is left without a current statusupdate beyond the defined response window.
- Coordinate with L2 engineersduring escalations — provide clear handover notes including timeline, alertcontext, initial diagnostics, and business impact assessment.
- Follow the priority matrixstrictly: P1 (15 min), P2 (30 min), P3 (4 hr), P4 (8 hr) response andescalation thresholds.
Platform & Service HealthChecks
- Execute scheduled shift healthchecks across all managed platforms — Azure resources, on-premises servers,network devices, security appliances, and employee services.
- Verify availability andperformance of core services: DNS, DHCP, NTP, Active Directory, and M365platform components.
- Monitor security platformdashboards (firewalls, EDR, proxy services) for health status and alert flags;escalate anomalies per defined procedures.
- Review Azure Cost Managementdashboards for unusual consumption spikes and flag to the lead for review.
Routine Operations &Maintenance Support
- Execute scheduled batch jobs,backup verifications, replication checks, and housekeeping tasks as per theoperational calendar.
- Support L2 and Specialistengineers during planned maintenance windows, patch cycles, and changeactivities — providing monitoring coverage and rollback readiness.
- Validate post-changeinfrastructure health after every approved change; raise a flag immediately ifanomalies are detected post-implementation.
Documentation & KnowledgeManagement
- Maintain precise shift handoverreports — open tickets, ongoing incidents, recent changes, and watch-items forthe next shift.
- Contribute to the knowledge baseby documenting recurring alert patterns, resolution steps, and workarounds forL1-resolvable issues.
- Flag gaps in runbooks or SOPs tothe operations lead so that documentation is continuously improved.
Required Skills & Experience
- 2 – 5 years of experience in IToperations, infrastructure support, or cloud managed services.
- Working knowledge of MicrosoftAzure: Azure Portal navigation, VM status checks, resource monitoring, andbasic troubleshooting using Azure Monitor and Log Analytics.
- Hands-on familiarity withWindows Server (2016/2019/2022) and Linux (RHEL/Ubuntu) — service management,log file reading, process monitoring, and basic fault diagnosis.
- Understanding of hybridinfrastructure models — on-premises datacenter integrated with Azure cloud viaExpressRoute or VPN.
- Basic familiarity with storageconcepts: disk performance thresholds, capacity monitoring, backup job status,and replication health checks.
Monitoring & ObservabilityTools
- Experience reading andinterpreting Prometheus metrics and Grafana dashboards — understanding panelthresholds, alert states, and trend data.
- Familiarity with Azure Monitoralerts and Log Analytics at a basic level — understanding alert rules, severitylevels, and affected resources.
- Ability to read sFlow or NetFlowtraffic dashboards for high-level network health assessment.
- APM dashboard familiarity —understanding response time trends, error rates, and service health indicators from tools such as Dynatrace, AppDynamics, or Azure Application Insights.
- Good understanding of corenetworking concepts: TCP/IP, DNS, DHCP, NTP, VLANs, and basic routing.
- Practical diagnostic skills:ping, traceroute, nslookup, netstat — to perform first-level connectivitychecks and provide meaningful diagnostics to L2.
- Familiarity with VPN tunnelhealth monitoring and firewall status dashboards at an operational level.
ITSM & Process Discipline
- Experience working with ITSMplatforms (ServiceNow, Zoho Desk, Freshservice, or equivalent) for incidentlogging, ticket updates, and escalation workflows.
- Understanding of ITIL incidentmanagement concepts — priority, impact, urgency, escalation paths, and SLAtracking.
- Ability to follow runbooks andSOPs precisely and consistently, including under pressure during major incidentscenarios.
- Clear and concise writtencommunication for incident tickets, shift handover notes, and escalationsummaries.
Soft Skills & Work Style
- Comfortable working in a 24x7rotational shift environment including night shifts, weekends, and publicholidays.
- High attention to detail —accurate logging, precise documentation, and consistent process adherence.
- Team-oriented with a proactiveattitude toward learning and improving operational knowledge.
- Ability to stay calm andsystematic during high-pressure P1/P2 incident scenarios.
Certification (Preferred)
Domain
- AZ-900 (AzureFundamentals)|AZ-104 (Administrator — working towards)
Service Management
Networking
- CompTIA Network+|Cisco CCNA (advantageous)
Monitoring
- Grafana Certified Associate(advantageous)