Site Reliability Engineer – Digital Banking
Abu Dhabi, United Arab Emirates | IT and Services | Full‑time | Minimum 5 years experience
Estimated salary range: AED 26,000–40,000 per month (confirm with employer).
Overview
We are looking for an experienced SRE or DevOps professional with strong expertise in observability, AIOps, cloud-native infrastructure, Kubernetes, Terraform, incident management, and high‑availability digital banking platforms. The role focuses on improving reliability, reducing operational risk, automating production workflows, and supporting seamless customer experiences in regulated banking environments.
Role Context
The Site Reliability Engineer will help maintain stability, scalability, security, and performance of business‑critical banking services. The role emphasizes building reliable operating models through SLIs, SLOs, error budgets, observability, incident response, automation, and AI‑driven reliability improvements.
Key Responsibilities
- Define, implement, and track SLIs, SLOs, error budgets, and reliability targets for critical digital banking services.
- Build practical observability across metrics, logs, traces, dashboards, and alerts using Dynatrace, Prometheus, Grafana, ELK stack, and related tools.
- Use Dynatrace Davis AI or equivalent AIOps platforms to identify anomalies, predict failures, reduce alert noise, and improve proactive incident prevention.
- Lead incident response activities including on-call triage, impact assessment, technical coordination, escalation handling, root-cause analysis, and blameless postmortems.
- Improve production readiness through release reviews, rollback planning, canary deployments, blue-green deployments, and safe rollout strategies.
- Support microservices-based platforms by improving service resilience, scalability, dependency handling, and inter-service reliability.
- Conduct capacity planning, performance tuning, resilience testing, and cost optimization for high-traffic production systems.
- Automate operational toil through runbooks, health checks, remediation scripts, self-healing workflows, and repeatable support processes.
- Work with DevOps teams to embed reliability gates, validation checks, monitoring standards, and deployment safeguards into CI/CD pipelines.
- Support CI/CD environments using GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps, or similar platforms.
- Own and improve observability and AIOps capabilities to support predictive alerting, intelligent automation, and faster issue detection.
- Maintain operational documentation, playbooks, incident records, troubleshooting procedures, and environment standards.
- Ensure operational practices align with internal controls, security expectations, compliance requirements, and banking regulatory standards.
- Analyze availability, performance, reliability, incident, and cost data to identify improvement opportunities.
- Provide escalation support and reliability guidance during major incidents affecting critical production systems.
Ideal Profile
- Minimum 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or production reliability roles.
- Experience building, managing, or supporting large-scale high-availability systems in banking, fintech, e-commerce, or data-intensive digital ecosystems.
- Bachelor’s degree in Computer Science or equivalent technical experience.
- Strong hands‑on experience with Linux environments, system diagnostics, and performance troubleshooting.
- Proven expertise in Terraform, Infrastructure as Code, automated infrastructure provisioning, and configuration management practices.
- Strong working knowledge of Kubernetes, container orchestration, microservices environments, and cloud-native architecture.
- Hands‑on experience with AWS is preferred; exposure to Azure or GCP is an advantage.
- Deep knowledge of Dynatrace, Davis AI, Prometheus, Grafana, ELK stack, observability practices, and AIOps-led monitoring.
- Experience implementing AI or ML-driven reliability solutions such as anomaly detection, predictive alerting, intelligent monitoring, or automated remediation.
- Practical understanding of CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps, or equivalent tools.
- Experience with Kafka, RabbitMQ, Redis, Aurora, RDS, and distributed system components.
- Strong scripting or programming skills in Python, Bash, Go, or similar languages.
- Strong analytical ability with confidence troubleshooting complex production systems and identifying root causes.
- Calm, structured, and clear communicator who can coordinate effectively during high-pressure incidents.
- Proactive mindset with interest in preventive improvements, automation, observability, and continuous reliability engineering.
- Comfortable working with distributed teams in a regulated, fast-paced digital banking environment.
Skills Set
Site reliability engineering, digital banking reliability, DevOps, AIOps, Dynatrace, Davis AI, Prometheus, Grafana, ELK stack, observability, SLIs, SLOs, error budgets, incident management, root-cause analysis, blameless postmortems, Terraform, Infrastructure as Code, Kubernetes, container orchestration, microservices, AWS, Azure, GCP, Linux, performance troubleshooting, GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps, CI/CD reliability gates, Kafka, RabbitMQ, Redis, Aurora, RDS, Python, Bash, Go, anomaly detection, predictive alerting, runbook automation, self-healing workflows, canary deployments, blue-green deployments, capacity planning, resilience testing, regulatory compliance support, production support.
Why Join Us
This opportunity is ideal for a Site Reliability Engineer who wants to work on critical digital banking systems where uptime, automation, observability, and security directly impact business. The role offers strong exposure to AIOps, Dynatrace Davis AI, Kubernetes, Terraform, cloud infrastructure, microservices reliability, and high-pressure incident management. It is a valuable position for an engineer who enjoys building resilient platforms, improving operational maturity, and supporting digital banking innovation in Abu Dhabi.