Senior System Specialist (L3) – Kochi, India (Onsite)
Permanent – Full Time
About the Role
The IT Operations Engineer at L3 role is a senior, technical role responsible for maintaining stability, high availability, and performance of Cyncly's critical production infrastructure including all the on-prem and cloud hosting. This role serves as the 3rd escalation point (Level 3) for complex operational incidents. The Engineer will utilize deep systems knowledge to proactively identify and resolve issues, manage monitoring systems, and lead maintenance projects. The Engineer will strictly adhere to ITIL processes, leveraging Freshservice as the core platform for Incident, Problem, and Change Management. The ideal candidate brings deep domain expertise across Microsoft technologies including Active Directory, Azure AD (Entra ID), and Microsoft Intune, alongside strong hands‑on experience with virtualization platforms, enterprise backup and disaster recovery, certificate lifecycle management, and cloud service models (IaaS, PaaS, and SaaS) PowerShell Scripting.
Key Responsibilities
- Serve as the primary Level 3 escalation point for complex and recurring incidents affecting hosted systems, services, servers, and network devices.
- Perform proactive and reactive health checks, diagnostics, and high-level troubleshooting to ensure continuous availability and performance of the production environment.
- Conduct Root Cause Analysis (RCA) for major incidents, developing and implementing permanent remediation actions to prevent recurrence.
- Manage incident and problem resolution workflows within the FreshService ITSM platform, ensuring clear communication and accurate documentation.
- Adhere to and promote ITIL-aligned processes, specifically in Incident, Problem, and Change Management.
Advanced System Administration
- Provide expert system administration and maintenance for enterprise hosting at on‑prem and cloud.
- Hands‑on experience in troubleshooting, patch administration, installation, and remote administration of systems and network.
- Demonstrate extensive Active Directory expertise including design, administration, and troubleshooting of AD DS, DNS, DHCP, Group Policy Objects (GPO), Organizational Units (OU), trust relationships, domain controller management, and AD replication.
- Administer and maintain Azure Active Directory (Entra ID) including hybrid identity configurations, Azure AD Connect, Conditional Access policies, Privileged Identity Management (PIM), and integration with enterprise applications via SAML and OAuth.
- Manage Microsoft Intune and endpoint management at an expert level, including device compliance policies, configuration profiles, application deployment, Windows Autopilot provisioning, and co‑management with SCCM/Configuration Manager.
- Administer and maintain enterprise PKI and certificate lifecycle management, including issuing, renewing, revoking, and monitoring SSL/TLS certificates across internal and public‑facing systems, and managing certificate authority (CA) infrastructure.
- Manage enterprise backup and disaster recovery systems, including configuration, scheduling, monitoring, and testing of backup jobs; ensure RPO and RTO objectives are met and maintain documented DR runbooks.
- High focus on automation and AI (infrastructure as code, scripting, etc.)
- Configure, maintain, and optimize core infrastructure services, including Proxy servers and web servers (e.g., Apache).
- Manage the M365, Azure and other core service admin centres.
Virtualization and Cloud Infrastructure
- Design, deploy, and administer virtualization platforms including VMware vSphere/ESXi and Microsoft Hyper‑V, covering VM lifecycle management, resource allocation, host configuration, clustering, and performance tuning.
- Manage VMware vCenter environments including host clusters, datastores, vSAN, distributed virtual switches, and snapshot management.
Monitoring and Infrastructure Maintenance
- Design, configure, and maintain enterprise monitoring solutions, specifically Zabbix, ensuring comprehensive coverage of servers, network devices, and critical services.
- Lead technical improvement and maintenance projects, including OS and package patching, system upgrades, and capacity expansion planning.
- Provide technical support and knowledge transfer to L1 and L2 Service Desk/Operations teams.
Required Qualifications
- Bachelor’s degree in computer science, Information Systems, or a similar relevant technical degree.
- Minimum of 10+ years of experience in a high‑availability IT Operations or Infrastructure role, with a strong focus on systems and networking.
- Windows and Linux Expertise: Deep, hands‑on administrative experience in Windows and Linux OS variants, including advanced troubleshooting and system hardening.
- Extensive domain expertise in Active Directory (AD DS), Azure Active Directory / Entra ID, and Microsoft Intune, including hybrid identity architecture and modern endpoint management.
- Proven experience managing enterprise virtualization platforms – VMware vSphere/ESXi and/or Microsoft Hyper‑V – in production environments.
- Demonstrated experience designing and maintaining enterprise backup and disaster recovery solutions, with knowledge of backup tools, replication strategies, and DR testing.
- Hands‑on experience managing certificate infrastructure including PKI, internal and public SSL/TLS certificates, and certificate authority management.
- Practical, working knowledge of IaaS, PaaS, and SaaS cloud service models, with strong Azure platform experience.
- Certifications: Microsoft Certifications (e.g., AZ‑104, MD‑102, SC‑300, MS‑102), VMware (VCP), CISCO certifications, and RHCSA/RHCE are highly desired.
- Expertise in patching enterprise systems and packages.
- Proven experience installing and configuring enterprise monitoring tools such as Zabbix.
- Strong understanding of network concepts, protocols, and the ability to monitor and troubleshoot network devices.
Desired Skills
- Experience with scripting languages (e.g., Bash, Python, or PowerShell) for automation, infrastructure‑as‑code, and operational efficiency gains.
- Advanced Active Directory experience covering multi‑domain and multi‑forest architectures, AD Federation Services (ADFS), AD sites and services optimization, and schema‑level administration.
- Experience with Azure AD (Entra ID) advanced capabilities including Entitlement Management, Access Reviews, Identity Protection, and Privileged Identity Management (PIM) beyond basic administration.
- Familiarity with virtualization platforms (e.g., VMware, Hyper‑V).
- Knowledge of backup and disaster recovery systems.
- Familiarity with IT service management (ITSM) frameworks, such as ITIL.
Core Competency Requirements
- Strong problem‑solving and analytical skills, with the ability to handle complex, high‑pressure situations.
- Excellent written and verbal communication skills, suitable for detailed technical documentation and stakeholder updates.
- Must be a self‑starter, highly energetic, and a collaborative team player.
- Ability to work across cross‑functional teams spanning infrastructure, security, cloud, and service desk disciplines.
- Demonstrated ability to take ownership of complex problems end‑to‑end, from initial escalation through to permanent resolution and documentation.