Site Reliability Engineer - Digital Banking

DiceTek UAE

Al Ruways Industrial City

On-site

AED 290,160 - 446,400

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

DiceTek UAE is seeking a Site Reliability Engineer for digital banking in Abu Dhabi. The role is full-time, requiring a minimum of 5 years' experience in SRE or DevOps, focusing on cloud infrastructure and automation.

The successful candidate will work on maintaining stability and performance of critical banking services, along with implementing observability practices and leading incident response.

Exposure to AIOps and tools like Kubernetes and Terraform is essential, offering an exciting opportunity to impact digital banking innovation.

Qualifications

  • Minimum 5 years of experience in Site Reliability Engineering, DevOps, or related roles.
  • Strong hands-on experience with Linux environments and system diagnostics.
  • Experience with AWS preferred; exposure to Azure or GCP is an advantage.

Responsibilities

  • Define and implement SLIs, SLOs, error budgets for digital banking services.
  • Build observability across metrics, logs, traces using Dynatrace and Grafana.
  • Lead incident response activities including triage and root-cause analysis.

Skills

Site reliability engineering
DevOps
AIOps
Terraform
Kubernetes
Incident management
Performance troubleshooting
Python
Linux

Education

Bachelor's degree in Computer Science or equivalent technical experience

Tools

Dynatrace
Prometheus
Grafana
ELK stack
GitHub Actions
Jenkins
GitLab CI/CD
Azure DevOps

Job description

Site Reliability Engineer – Digital Banking

Abu Dhabi, United Arab Emirates | IT and Services | Full‑time | Minimum 5 years experience

Estimated salary range: AED 26,000–40,000 per month (confirm with employer).

Overview

We are looking for an experienced SRE or DevOps professional with strong expertise in observability, AIOps, cloud-native infrastructure, Kubernetes, Terraform, incident management, and high‑availability digital banking platforms. The role focuses on improving reliability, reducing operational risk, automating production workflows, and supporting seamless customer experiences in regulated banking environments.

Role Context

The Site Reliability Engineer will help maintain stability, scalability, security, and performance of business‑critical banking services. The role emphasizes building reliable operating models through SLIs, SLOs, error budgets, observability, incident response, automation, and AI‑driven reliability improvements.

Key Responsibilities
  • Define, implement, and track SLIs, SLOs, error budgets, and reliability targets for critical digital banking services.
  • Build practical observability across metrics, logs, traces, dashboards, and alerts using Dynatrace, Prometheus, Grafana, ELK stack, and related tools.
  • Use Dynatrace Davis AI or equivalent AIOps platforms to identify anomalies, predict failures, reduce alert noise, and improve proactive incident prevention.
  • Lead incident response activities including on-call triage, impact assessment, technical coordination, escalation handling, root-cause analysis, and blameless postmortems.
  • Improve production readiness through release reviews, rollback planning, canary deployments, blue-green deployments, and safe rollout strategies.
  • Support microservices-based platforms by improving service resilience, scalability, dependency handling, and inter-service reliability.
  • Conduct capacity planning, performance tuning, resilience testing, and cost optimization for high-traffic production systems.
  • Automate operational toil through runbooks, health checks, remediation scripts, self-healing workflows, and repeatable support processes.
  • Work with DevOps teams to embed reliability gates, validation checks, monitoring standards, and deployment safeguards into CI/CD pipelines.
  • Support CI/CD environments using GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps, or similar platforms.
  • Own and improve observability and AIOps capabilities to support predictive alerting, intelligent automation, and faster issue detection.
  • Maintain operational documentation, playbooks, incident records, troubleshooting procedures, and environment standards.
  • Ensure operational practices align with internal controls, security expectations, compliance requirements, and banking regulatory standards.
  • Analyze availability, performance, reliability, incident, and cost data to identify improvement opportunities.
  • Provide escalation support and reliability guidance during major incidents affecting critical production systems.
Ideal Profile
  • Minimum 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud operations, or production reliability roles.
  • Experience building, managing, or supporting large-scale high-availability systems in banking, fintech, e-commerce, or data-intensive digital ecosystems.
  • Bachelor’s degree in Computer Science or equivalent technical experience.
  • Strong hands‑on experience with Linux environments, system diagnostics, and performance troubleshooting.
  • Proven expertise in Terraform, Infrastructure as Code, automated infrastructure provisioning, and configuration management practices.
  • Strong working knowledge of Kubernetes, container orchestration, microservices environments, and cloud-native architecture.
  • Hands‑on experience with AWS is preferred; exposure to Azure or GCP is an advantage.
  • Deep knowledge of Dynatrace, Davis AI, Prometheus, Grafana, ELK stack, observability practices, and AIOps-led monitoring.
  • Experience implementing AI or ML-driven reliability solutions such as anomaly detection, predictive alerting, intelligent monitoring, or automated remediation.
  • Practical understanding of CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps, or equivalent tools.
  • Experience with Kafka, RabbitMQ, Redis, Aurora, RDS, and distributed system components.
  • Strong scripting or programming skills in Python, Bash, Go, or similar languages.
  • Strong analytical ability with confidence troubleshooting complex production systems and identifying root causes.
  • Calm, structured, and clear communicator who can coordinate effectively during high-pressure incidents.
  • Proactive mindset with interest in preventive improvements, automation, observability, and continuous reliability engineering.
  • Comfortable working with distributed teams in a regulated, fast-paced digital banking environment.
Skills Set

Site reliability engineering, digital banking reliability, DevOps, AIOps, Dynatrace, Davis AI, Prometheus, Grafana, ELK stack, observability, SLIs, SLOs, error budgets, incident management, root-cause analysis, blameless postmortems, Terraform, Infrastructure as Code, Kubernetes, container orchestration, microservices, AWS, Azure, GCP, Linux, performance troubleshooting, GitHub Actions, Jenkins, GitLab CI/CD, Azure DevOps, CI/CD reliability gates, Kafka, RabbitMQ, Redis, Aurora, RDS, Python, Bash, Go, anomaly detection, predictive alerting, runbook automation, self-healing workflows, canary deployments, blue-green deployments, capacity planning, resilience testing, regulatory compliance support, production support.

Why Join Us

This opportunity is ideal for a Site Reliability Engineer who wants to work on critical digital banking systems where uptime, automation, observability, and security directly impact business. The role offers strong exposure to AIOps, Dynatrace Davis AI, Kubernetes, Terraform, cloud infrastructure, microservices reliability, and high-pressure incident management. It is a valuable position for an engineer who enjoys building resilient platforms, improving operational maturity, and supporting digital banking innovation in Abu Dhabi.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - AIOps
Site Reliability Engineer - AIOps

DiceTek UAE • Al Ruways Industrial City

On-site
Senior SRE: AI-Driven Reliability for Digital Banking
Senior SRE: AI-Driven Reliability for Digital Banking

DiceTek UAE • Al Ruways Industrial City

On-site
Senior DevOps / Site Reliability Engineer (SRE)
Senior DevOps / Site Reliability Engineer (SRE)

Stellar Technologies • Abu Dhabi

On-site
AED 360,000 - 540,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Dubai

On-site
AED 200,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

Deeplight • United Arab Emirates

On-site
AED 350,000 - 550,000
Competitive salary
Comprehensive personal health 보험
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AW Connect • Dubai

On-site
AED 420,000 - 660,000
Technology Service Lead - Digital Banking
Technology Service Lead - Digital Banking

First Abu Dhabi Bank • Al Ruways Industrial City

On-site
Senior SRE - AIOps & Cloud Reliability Engineer
Senior SRE - AIOps & Cloud Reliability Engineer

DiceTek UAE • Al Ruways Industrial City

On-site
Head - Developer Experience and DevSecOps
Head - Developer Experience and DevSecOps

RAKBANK • Dubai

On-site
AED 60,000 - 85,000
Innovation-focused environment
Leadership in AI-native banking transformation