Location: Abu Dhabi, UAE
Employment: 12-months initial (contract to perm)
We are seeking a hands-on Senior DevOps / Site Reliability Engineer to build and operate the delivery and runtime foundations for a portfolio of modern enterprise applications, workflow platforms and AI-enabled products.
This is not primarily an infrastructure-administration role. You will work directly with software engineers to create reliable, secure and automated paths from code to production.
You will own deployment automation, runtime reliability, observability, infrastructure-as-code and operational readiness across applications that integrate with critical enterprise systems.
What You Will Own
Platform & Infrastructure
- Build and maintain cloud infrastructure using Infrastructure as Code.
- Design secure, repeatable environments across development, test, staging and production.
- Manage containerized workloads and Kubernetes-based deployments where appropriate.
- Define standard application deployment patterns for backend, frontend and AI services.
- Implement secure secrets and configuration management.
- Support network, identity and connectivity requirements for enterprise integrations.
CI/CD & Developer Productivity
- Build automated CI/CD pipelines.
- Standardize build, test, security scanning and deployment processes.
- Automate environment provisioning and configuration.
- Reduce manual deployment steps and production configuration drift.
- Work closely with engineering teams to improve release frequency and reliability.
Reliability & Observability
- Establish logging, metrics, tracing and alerting.
- Define service-level indicators and operational thresholds.
- Build dashboards for system health and application performance.
- Implement incident-response and production-support practices.
- Design for graceful degradation, retries, failover and recovery.
- Lead root-cause analysis of production incidents.
Security & Operational Controls
- Implement least-privilege access and secure deployment patterns.
- Support auditability of infrastructure and production changes.
- Integrate security checks into delivery pipelines.
- Work with security and infrastructure teams to meet enterprise control requirements.
Resilience
- Support business continuity and disaster-recovery design.
- Define backup, restore and recovery procedures.
- Test operational recovery rather than relying solely on documented plans.
Required Experience
- 6+ years in DevOps, SRE, platform engineering or cloud infrastructure.
- Strong production experience with Azure, AWS or GCP; Azure strongly preferred.
- Docker and Kubernetes.
- Infrastructure as Code using Terraform, Bicep, Pulumi or equivalent.
- CI/CD using Azure DevOps, GitHub Actions, GitLab CI or similar.
- Strong Linux and networking fundamentals.
- Observability tooling and distributed-system troubleshooting.
- Secure secrets, identity and access-management patterns.
- Production incident-management experience.
- Scripting/programming capability in Python, Go, Bash or equivalent.
Strong Advantage
- Azure Kubernetes Service.
- Azure Service Bus, API Management, Key Vault and related Azure services.
- Enterprise integration platforms.
- SAP-connected environments.
- AI/LLM application deployment.
- Regulated or government environments.
- High-availability and disaster-recovery architecture.