Job Title - Cloud DevOps Lead – Specialist - ACS SONG
Management Level: Level 9 - Specialist
Location: Kochi
Must-have skills: Cloud operations leadership; installation, configuration and management of Linux systems; public cloud administration; incident and problem management
Good-to-have skills: Terraform/CloudFormation, Docker/Kubernetes, monitoring and observability tools, scripting/automation, AI adoption and AI-assisted operations, JIRA/Confluence
Experience: 5–8 years of relevant experience, including at least 2 years in a technical leadership or team‑lead role
Educational Qualification: Graduation
Job Summary
As a Cloud DevOps Lead, you will provide technical leadership for an e-commerce platform running across on‑premises and cloud infrastructure. You will own platform reliability, operational readiness, and technical delivery; lead the resolution of complex incidents and problems; govern production changes; and drive automation, observability, security, performance, and cost improvements. You will guide engineers, coordinate with application, infrastructure, security, and business stakeholders, and act as a trusted, client‑facing technical advisor by identifying improvement opportunities and presenting clear, value‑focused proposals that strengthen reliability, efficiency, security, and cost effectiveness.
Roles And Responsibilities
- Lead day‑to‑day cloud operations and provide technical direction, task prioritization, coaching, and escalation support to the operations team.
- Build trusted client relationships through regular technical discussions, service reviews, and advisory sessions; understand business priorities and translate them into practical cloud and operational improvement opportunities.
- Develop and present evidence‑based technical improvement proposals, including the current‑state assessment, recommended solution, expected business and operational benefits, risks, effort, cost considerations, and implementation roadmap; incorporate client feedback and support decisions through execution.
- Own platform availability, reliability, performance, capacity, security, and operational readiness across on‑premises and public‑cloud environments.
- Lead major incident response, coordinate technical recovery, communicate status to stakeholders, and ensure timely root‑cause analysis and preventive actions.
- Govern problem, change, release, service‑request, and configuration‑management activities in line with service‑level agreements and operational controls.
- Review system architecture, deployment designs, operational procedures, and technical changes to ensure scalability, resilience, supportability, and compliance.
- Define and improve monitoring, alerting, logging, dashboards, runbooks, and service health indicators to enable proactive issue detection and resolution.
- Drive automation of provisioning, deployment, maintenance, recovery, and routine support activities using infrastructure‑as‑code and scripting practices.
- Analyze operational trends, recurring incidents, capacity risks, and service metrics; maintain an improvement backlog and track actions to closure.
- Partner with application, DevOps, cloud, network, database, security, and vendor teams to resolve cross‑platform issues and deliver technical improvements.
- Maintain technical documentation, knowledge articles, disaster‑recovery procedures, and audit evidence; conduct knowledge‑sharing and readiness reviews.
- Support effort estimation, technical planning, resource coordination, and stakeholder reporting for operational and transformation initiatives.
- Participate in the on‑call rotation and ensure effective shift handovers, escalation paths, and operational coverage.