Overview
Site Reliability Engineer - AIOps is an IT and services opportunity for an experienced SRE or DevOps professional with strong expertise in high‑availability systems, cloud infrastructure, Kubernetes, Terraform, observability and AI‑driven reliability engineering. The role supports the reliability, scalability, performance and operational stability of mission‑critical digital systems in Abu Dhabi, United Arab Emirates.
Gender: Any Candidate Nationality: Any
Job Details
Country: United Arab Emirates
City: Abu Dhabi
Industry: IT and Services
Function: Information Technology
Salary: 25,000 – 38,000 (USD) per month
Prime Responsibilities
- Build, manage and improve large‑scale, high‑availability systems across cloud‑native and data‑intensive digital platforms.
- Design and maintain Infrastructure as Code using Terraform and modern IaC methodologies.
- Operate and optimise Kubernetes‑based container orchestration environments for microservices platforms.
- Support cloud infrastructure on AWS, with additional exposure to Azure or GCP considered an advantage.
- Monitor Linux environments, troubleshoot performance issues and improve system stability across production workloads.
- Implement observability practices using Dynatrace, Davis AI, Prometheus, Grafana, ELK stack, logs, metrics, traces, alerts and dashboards.
- Apply AI and ML‑driven reliability practices, including AIOps, anomaly detection, predictive alerting, automated remediation and intelligent incident insights.
- Improve CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps or similar delivery platforms.
- Support production systems that use Kafka, RabbitMQ, Redis, Aurora, RDS and related distributed technology components.
- Develop automation scripts and operational tools using Python, Bash, Go or similar programming languages.
- Lead troubleshooting during high‑impact incidents, coordinate technical response and support root‑cause analysis.
- Collaborate with SRE, DevOps, application, database, security and infrastructure teams to improve platform reliability.
- Create preventive reliability improvements based on incident trends, system metrics, capacity risks and recurring failure patterns.
- Document operational procedures, incident findings, monitoring standards, deployment practices and reliability improvements.
- Promote SRE best practices across cloud‑native, microservices, regulated technology and digital banking environments.
Ideal Profile
- Minimum 5 years of experience in SRE, DevOps, platform engineering, cloud operations or production reliability roles.
- Experience building and managing large‑scale, high‑availability systems in banking, fintech, e‑commerce or other data‑intensive digital environments.
- Bachelor’s degree in Computer Science or equivalent technical experience.
- Strong Linux administration, performance troubleshooting and production system debugging skills.
- Hands‑on expertise with Terraform, Infrastructure as Code and automated infrastructure provisioning.
- Strong working knowledge of Kubernetes, container orchestration, microservices and cloud‑native architecture.
- Practical experience with AWS cloud services is preferred; Azure or GCP exposure will be beneficial.
- Deep knowledge of Dynatrace, AIOps, Davis AI, Prometheus, Grafana and ELK stack.
- Experience implementing AI or ML‑driven reliability solutions such as anomaly detection, predictive alerting, intelligent monitoring and automation.
- Good understanding of CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps or similar tools.
- Experience with Kafka, RabbitMQ, Redis, Aurora, RDS and distributed application components.
- Strong scripting or programming ability in Python, Bash, Go or related languages.
- Calm and structured communicator who can coordinate during incidents and work well under pressure.
- Organised, meticulous, proactive and comfortable working with distributed teams in a fast‑paced regulated environment.
- Strong interest in automation, observability, cloud‑native platforms and continuous reliability improvement.
Skills Set
- Site reliability engineering
- DevOps
- AIOps
- Dynatrace
- Davis AI
- Prometheus
- Grafana
- ELK stack
- Terraform
- Infrastructure as Code
- Kubernetes
- Container orchestration
- AWS
- Azure
- GCP
- Linux
- Microservices
- CI/CD
- GitHub Actions
- Jenkins
- GitLab CI/CD
- Azure DevOps
- Kafka
- RabbitMQ
- Redis
- Aurora
- RDS
- Python
- Bash
- Go
- Anomaly detection
- Predictive alerting
- Observability
- Monitoring
- Logging
- Incident management
- Root‑cause analysis
- Performance troubleshooting
- Production reliability
- Cloud‑native systems
Why Join Us
This opportunity is well suited for an SRE or DevOps engineer who wants to work on high‑availability platforms, AI‑driven reliability, cloud‑native infrastructure and digital banking technology in Abu Dhabi. The role provides strong exposure to Kubernetes, Terraform, AWS, observability platforms, AIOps, automation and large‑scale production systems where reliability engineering directly supports business continuity, customer trust and digital innovation.
About the Company
Dicetek LLC is a technology services and consulting company supporting clients across the UAE and the wider region with skilled professionals in IT, cloud engineering, DevOps, cybersecurity, data, software development and digital transformation. The company works with organisations that require reliable technology talent to strengthen enterprise systems, deliver scalable platforms and support long‑term business growth through modern digital solutions.