About the Role
We are looking for a DevOps Engineer to join our Cloud Service Desk (CSD) team, acting as the primary bridge between Cloud Infrastructure, Engineering, and IT Operations. This role combines hands-on DevOps engineering with Cloud Service Desk responsibilities, supporting business-critical applications and cloud platforms across a hybrid AWS and Azure environment.
The successful candidate will primarily work in a Cloud Service Desk shift model, providing operational support, incident response, infrastructure management, and deployment coordination for both Microsoft and Java-based application stacks. You will play a key role in ensuring platform availability, operational excellence, and seamless collaboration between engineering and infrastructure teams.
Role & responsibilities
DevOps & Cloud Platform Support
- Act as the primary liaison between Engineering and Cloud Infrastructure teams, ensuring smooth deployment, support, and operational processes.
- Manage and support workloads across AWS and Azure cloud environments, including:
- AWS ECS, EKS, S3, CloudFront, EC2
- Azure IaaS, PaaS, and associated cloud services
- Support Microsoft (.NET) and Java application environments hosted across Windows and Linux platforms.
- Configure, maintain, and troubleshoot API Gateways, web servers, and application servers.
- Perform IAM access management, security policy administration, and lifecycle management activities.
- Investigate and resolve infrastructure, platform, application, and deployment issues across multi-cloud environments.
- Recommend and implement improvements to platform reliability, performance, security, and cost optimisation.
- Support CI/CD pipelines and change deployments within production and non-production environments.
- Participate in incident, change, problem, and release management activities.
Cloud Service Desk (CSD) Operations
- Primarily work within the Cloud Service Desk (CSD) shift structure, providing operational support for cloud-hosted services and applications.
- Act as an escalation point for 1st and 2nd Line Support teams on cloud, infrastructure, and application-related incidents.
- Monitor cloud services, application health, infrastructure alerts, and operational dashboards.
- Manage support tickets, requests, incidents, service restoration activities, and change records in accordance with ITIL best practices.
- Perform initial triage, root cause analysis, troubleshooting, and service recovery activities.
- Coordinate with Engineering, Infrastructure, Security, and Vendor teams during major incidents and production outages.
- Ensure adherence to SLA and KPI targets through timely resolution and effective communication.
- Participate in shift handovers and maintain operational documentation, runbooks, and knowledge base articles.
- Support maintenance activities, cloud platform upgrades, patching, and scheduled operational tasks.
- Contribute to continuous service improvement initiatives and operational automation.
Preferred candidate profile
- 3-5 years of experience in DevOps, Cloud Infrastructure, or Site Reliability roles or Cloud Operations roles
- Hands-on experience with AWS services (ECS, EKS, S3, CloudFront) and Azure
- Experience managing Windows-based EC2 instances alongside Linux environments
- Familiarity with both Microsoft/.NET and Java application stacks
- Experience with API gateways, web servers (e.g., IIS, Nginx, Apache), and application servers
- Working knowledge of IAM policies, access management, and lifecycle/retention policies
- Strong communication skills, comfortable working cross-functionally between engineering and infrastructure teams
- Flexibility for 24-hour shift work or out-of-hours support where required.
- Ability to support production cloud-native platforms and participate in incident, change, and release activities as required.
Nice to Have
- Experience with CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation, ARM/Bicep)
- Relevant AWS or Azure certifications