Site Reliability Engineer (SRE)

Dicetek LLC

Abu Dhabi

On-site

AED 180,000 - 280,000

Full time

30 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Dicetek LLC in Abu Dhabi is seeking an experienced SRE/DevOps engineer to define and own reliability metrics for our digital banking services. You will build observability with Dynatrace, Prometheus, Grafana, and ELK, lead incident management, and implement safe deployment strategies (canary/blue-green) while integrating reliability checks into CI/CD workflows.

You will collaborate across teams to ensure scalability, cost efficiency, and regulatory alignment, leveraging Terraform and cloud

Qualifications

  • 5+ years of experience in SRE or DevOps across banking or fintech domains.
  • Bachelor’s degree in Computer Science or equivalent.
  • Strong experience with Linux environments and performance troubleshooting.
  • Proficiency with Terraform and Infrastructure as Code.
  • Hands-on experience with Kubernetes and cloud providers (AWS preferred; Azure/GCP a plus).
  • Deep knowledge of Dynatrace, Prometheus, Grafana, and the ELK stack.
  • Experience implementing AI / ML-driven reliability or automation solutions (AIOps, anomaly detection, predictive alerting).
  • Practical understanding of CI / CD pipelines (GitHub Actions, Jenkins, GitLab CI / CD or Azure DevOps).

Responsibilities

  • Define and implement SLI/SLOs and error budgets for critical services.
  • Build observability with metrics, logs, traces, dashboards, and alerts.
  • Lead incident management from on-call triage to postmortems.
  • Implement safe deployment strategies including canaries and blue-green rollouts.
  • Collaborate with DevOps to embed reliability checks in CI/CD.
  • Optimize capacity, performance, and resilience of microservices.
  • Automate toil with runbooks, remediation scripts, and health checks.
  • Maintain documentation and compliance with security controls.

Skills

SRE
DevOps
Linux troubleshooting
CI/CD
Analytical mindset
Incident leadership

Education

Bachelor’s degree in Computer Science or equivalent

Tools

Dynatrace
Prometheus
Grafana
ELK stack
Terraform
Kubernetes
AWS
Azure
GCP
GitHub Actions
Jenkins
GitLab CI
Azure DevOps
Kafka
RabbitMQ
Redis
Aurora
RDS
Python
Bash
Go

Job description

  • Define and implement SLIs / SLOs and error budgets for business-critical digital banking services.
  • Build actionable observability (metrics, logs, traces, dashboards, and alerts) using Dynatrace, Prometheus, Grafana, and ELK, while reducing alert fatigue.
  • Leverage AI-driven insights and anomaly detection (Dynatrace Davis AI or equivalent AIOps platform) to proactively predict and resolve reliability issues before impact.
  • Lead incident management — from on-call triage and root-cause analysis to blameless postmortems with actionable follow-ups.
  • Improve deployment safety with robust rollout / rollback strategies, canary and blue-green deployments, and production readiness reviews.
  • Support and optimize microservices-based architectures, ensuring service reliability, scalability, and inter-service resilience.
  • Conduct capacity planning, performance tuning, and resilience testing, optimizing for both reliability and cost efficiency.
  • Automate operational toil — from runbooks and remediation scripts to proactive health checks and self-healing workflows.
  • Collaborate with DevOps to embed reliability gates and validations into CI / CD pipelines (GitHub Actions, Jenkins, GitLab CI / CD or Azure DevOps).
  • Own and evolve the observability and AIOps stack, driving intelligent automation and predictive alerting capabilities.
  • Maintain high-quality documentation, playbooks, and operational standards across environments.
  • Ensure operational compliance and security alignment with internal controls and regulatory standards.
  • Analyze system performance, availability, and cost data to continually optimize operations.
  • Provide reliability support and escalation guidance for critical production systems during major incidents.
What You Will Be Doing
  • Define and implement SLIs / SLOs and error budgets for business-critical digital banking services.
  • Build actionable observability (metrics, logs, traces, dashboards, and alerts) using Dynatrace, Prometheus, Grafana, and ELK, while reducing alert fatigue.
  • Leverage AI-driven insights and anomaly detection (Dynatrace Davis AI or equivalent AIOps platform) to proactively predict and resolve reliability issues before impact.
  • Lead incident management — from on-call triage and root-cause analysis to blameless postmortems with actionable follow-ups.
  • Improve deployment safety with robust rollout / rollback strategies, canary and blue-green deployments, and production readiness reviews.
  • Support and optimize microservices-based architectures, ensuring service reliability, scalability, and inter-service resilience.
  • Conduct capacity planning, performance tuning, and resilience testing, optimizing for both reliability and cost efficiency.
  • Automate operational toil — from runbooks and remediation scripts to proactive health checks and self-healing workflows.
  • Collaborate with DevOps to embed reliability gates and validations into CI / CD pipelines (GitHub Actions, Jenkins, GitLab CI / CD or Azure DevOps).
  • Own and evolve the observability and AIOps stack, driving intelligent automation and predictive alerting capabilities.
  • Maintain high-quality documentation, playbooks, and operational standards across environments.
  • Ensure operational compliance and security alignment with internal controls and regulatory standards.
  • Analyze system performance, availability, and cost data to continually optimize operations.
  • Provide reliability support and escalation guidance for critical production systems during major incidents.
Experience And Qualifications
  • 5+ years of experience in SRE or DevOps roles, building and managing large-scale, high-availability systems across banking, fintech, e-commerce, or other data-intensive digital ecosystems.
  • Bachelor’s degree in Computer Science or equivalent technical experience.
  • Strong experience with Linux environments and performance troubleshooting.
  • Proven expertise in Terraform and Infrastructure as Code (IaC) methodologies.
  • Proficiency with Kubernetes and container orchestration in microservices environments.
  • Hands-on experience with AWS (preferred); exposure to Azure or GCP is an advantage.
  • Deep knowledge of Dynatrace (AIOps, Davis AI), Prometheus, Grafana, and the ELK stack.
  • Experience implementing AI / ML-driven reliability or automation solutions (AIOps, anomaly detection, predictive alerting).
  • Practical understanding of CI / CD pipelines (GitHub Actions, Jenkins, GitLab CI / CD or Azure DevOps).
  • Experience with Kafka, RabbitMQ, Redis, Aurora, and RDS databases.
  • Strong scripting or programming skills in Python, Bash, or Go. The Ideal Candidate
  • Organized, structured, and meticulous in approach.
  • Experienced in cross-functional collaboration and working with distributed teams.
  • Strong analytical mindset with excellent troubleshooting skills for complex production systems.
  • Calm and composed communicator under pressure, capable of leading during high-impact incidents.
  • Proactive problem-solver who anticipates issues and drives preventive improvements.
  • Passionate about AI-driven automation, observability, and reliability engineering.
  • Continuously learning, keeping up-to-date with cloud-native, microservices, and SRE best practices.
  • A collaborative and adaptable team player who thrives in a fast-paced, regulated environment and is passionate about building reliable, scalable systems that empower digital banking innovation.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE (Site Reliability Engineer)
SRE (Site Reliability Engineer)

Dicetek LLC • Abu Dhabi

On-site
AED 180,000 - 300,000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Salt • Abu Dhabi

On-site
AED 480,000 - 720,000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Client of Salt • Abu Dhabi

On-site
AED 320,000 - 520,000
Senior DevOps / Site Reliability Engineer (SRE)
Senior DevOps / Site Reliability Engineer (SRE)

Stellar Technologies • Abu Dhabi

On-site
AED 360,000 - 540,000
Site Reliability Engineer (SRE) - Azure focus
Site Reliability Engineer (SRE) - Azure focus

Dicetek LLC • Dubai

On-site
AED 300,000 - 550,000
Lead SRE / Technology Operations (DevSecOps)
Lead SRE / Technology Operations (DevSecOps)

Tanqeeb • Dubai

On-site
AED 350,000 - 660,000
Lead SRE / Technology Operations (DevSecOps)
Lead SRE / Technology Operations (DevSecOps)

Client of FinTop Consulting • Dubai

On-site
AED 420,000 - 640,000
Senior DevOps / SRE Engineer
Senior DevOps / SRE Engineer

GSSTech Group • Dubai

On-site
AED 441,000 - 698,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

31 CONCEPT • United Arab Emirates

On-site
AED 300,000 - 460,000
Principal SRE
Principal SRE

Alpheya • Abu Dhabi

On-site
AED 900,000 - 1,200,000