Site Reliability Engineer (SRE) | Innovations Global | Abu Dhabi, UAE

Innovations Global

Abu Dhabi

On-site

AED 350,000 - 550,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Innovations Global is urgently seeking an experienced Site Reliability Engineer (SRE) to champion availability, scalability, and performance for banking platforms in Abu Dhabi on a 1-year renewable contract. You will bridge software engineering and operations, lead observability, automation, incident recovery, and enforce SRE principles across hybrid cloud environments.

The role requires 5+ years in SRE/DevOps, strong Linux/Unix, Docker, Kubernetes, and cloud experience.

Qualifications

  • 5+ years of SRE/DevOps or Linux/Unix experience.
  • Strong Linux/Unix administration and performance tuning.
  • Experience with AWS, Azure, or GCP.
  • Proficiency with Docker and Kubernetes.
  • Experience with monitoring/observability tools (Prometheus, Grafana, ELK, Datadog, Dynatrace).
  • Scripting in Python or Bash for automation.

Responsibilities

  • Ensure reliability, availability, and performance of banking apps and cloud infra.
  • Implement full-stack monitoring and real-time alerts.
  • Automate workflows with Python/Bash to reduce toil.
  • Lead incident management, RCAs, and permanent remediation.
  • Deploy and scale containers with Kubernetes and Docker.
  • Maintain Linux/Unix servers in Tier-compliant hosting.
  • Define and report SLA/SLO/SLI metrics.
  • Improve CI/CD pipelines for secure releases.

Skills

Linux/Unix
Docker
Kubernetes
Monitoring
Python/Bash
Cloud platforms
Networking
SRE/DevOps

Tools

Prometheus
Grafana
ELK
Datadog
Dynatrace

Job description

Position Summary

Innovations Global is urgently seeking an experienced, proactive Site Reliability Engineer (SRE) to champion the continuous availability, scalability, and performance of mission-critical banking and financial technology platforms in Abu Dhabi, United Arab Emirates. Operating in a full-time onsite capacity for a 1-year renewable contract, you will bridge the divide between software engineering and systems operations to ensure enterprise infrastructure resilience. With a minimum of five years of hands-on experience, you will spearhead observability, drive end-to-end automation, lead rapid incident recovery, and enforce stringent Site Reliability Engineering principles across hybrid cloud environments. This position provides an exceptional opportunity for a technical engineering professional to make a transformative impact on premier banking infrastructure in the UAE.

Detailed Job Description

As a Site Reliability Engineer (SRE) at Innovations Global supporting a premier enterprise banking environment in Abu Dhabi, you will hold operational accountability for the health, availability, performance, and efficiency of high-throughput transactional applications and distributed systems. You will collaborate closely with cross-functional software engineering teams, DevOps squads, database administrators, and cyber security teams to establish resilient deployment pipelines and maintain robust production ecosystems.

Your core technical mandate involves architecting and administering enterprise Linux and Unix servers, orchestrating microservices utilizing Docker and Kubernetes, and managing scalable workloads across leading cloud platforms (AWS, Azure, or GCP). You will implement proactive monitoring and observability frameworks, define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs), and eliminate operational toil through Python and Bash automation scripting. Additionally, you will direct incident response triage, execute root cause analysis (RCA), optimize CI/CD release workflows, and troubleshoot complex TCP/IP enterprise networking bottlenecks. This role requires rigorous diagnostic discipline, deep systems acumen, and the capability to maintain zero-downtime reliability within a fast-paced financial services environment.

Key Responsibilities
  • Ensure the maximum reliability, availability, performance, and operational efficiency of enterprise banking applications and underlying cloud infrastructure.
  • Implement, tune, and manage full-stack monitoring, telemetry, and observability platforms to capture proactive operational insights and real-time alerts.
  • Eliminate operational toil by designing, building, and maintaining automated workflows and operational scripts using Python, Bash, or Shell scripting.
  • Lead rapid incident management, triage system outages, conduct detailed root cause analysis (RCA), and implement permanent corrective remediations.
  • Deploy, configure, manage, and scale containerized application workloads utilizing Kubernetes clusters and Docker environments.
  • Administer, optimize, and maintain high-performance enterprise Linux and Unix server operating systems in Tier-compliant hosting environments.
  • Establish, monitor, and report on core reliability engineering metrics, including Service Level Agreements (SLAs), Service Level Objectives (SLOs), and Error Budgets.
  • Optimize and support automated CI/CD deployment pipelines, ensuring secure, reliable, and frictionless software releases into production.
Required Qualifications & Skills
  • Minimum 5+ years of dedicated professional experience in Site Reliability Engineering (SRE), DevOps, or Linux/Unix Systems Engineering.
  • Strong technical expertise in Linux and Unix system administration, operating system internals, kernel parameters, and performance tuning.
  • Hands-on expertise deploying, managing, and maintaining enterprise workloads on major public cloud platforms such as AWS, Microsoft Azure, or GCP.
  • Demonstrated technical proficiency with containerization and orchestration platforms, specifically Docker and Kubernetes.
  • Extensive experience implementing modern monitoring, logging, and observability tools (e.g., Prometheus, Grafana, ELK Stack, Datadog, or Dynatrace).
  • Proven proficiency in scripting and automation utilizing Python, Bash, or Shell scripting for operational workflows.
  • Solid understanding of core TCP/IP networking, routing, DNS, load balancing, SSL/TLS, and enterprise perimeter network security.
  • Demonstrated experience in incident management, blameless post-mortem investigations, and reliability engineering practices (SLA/SLO/SLI frameworks).
Nice-to-Have Skills
  • Prior hands-on engineering experience within the banking, financial services, fintech, or large-scale transactional enterprise domains.
  • Recognized professional certifications such as AWS Certified Solutions Architect, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
  • Familiarity with Infrastructure as Code (IaC) tooling including Terraform, Ansible, or CloudFormation.
  • Experience implementing GitOps deployment paradigms, ArgoCD, Jenkins, or GitLab CI/CD automated release pipelines.
  • Current physical residence in Abu Dhabi or elsewhere in the United Arab Emirates with immediate joining capability.
Application Information
  • Employer / Recruiting Agency: Innovations Global
  • Contact Person: S.P. Hariharan
  • Role: Site Reliability Engineer (SRE)
  • Work Location: Abu Dhabi, UAE (100% Onsite)
  • Job Type: Contract (1 year renewable project)
  • Notice Period: Immediate to 30 Days
  • Experience Required: 5+ years
  • Target Domain: Banking, Financial Services, Fintech, or Enterprise Applications
  • Application Email: sp.hariharan@innovationsglobal.com
Recruitment Pro Tip

Given that this role supports an enterprise banking environment in Abu Dhabi with an immediate-to-30-day onboarding window, highlight your direct financial services or high-concurrency transactional systems background on page one of your CV. Clearly articulate the specific observability stacks (Prometheus, Grafana, ELK) you have architected, quantify the repetitive toil you eliminated via Python or Bash scripting, and explicitly note your current location and notice period when emailing sp.hariharan@innovationsglobal.com.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE: Observability, Automation & Incident Response (Abu Dhabi)
SRE: Observability, Automation & Incident Response (Abu Dhabi)

Innovations Global • Abu Dhabi

On-site
AED 350,000 - 550,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Al Ain

On-site
AED 180,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether SRL • United Arab Emirates

On-site
AED 350,000 - 600,000
Fully remote
Global engineering org
Ownership over reliability
+2
Senior DevOps / Site Reliability Engineer (SRE)
Senior DevOps / Site Reliability Engineer (SRE)

Stellar Technologies • Abu Dhabi

On-site
AED 360,000 - 540,000
SRE (Site Reliability Engineer)
SRE (Site Reliability Engineer)

Dicetek LLC • Abu Dhabi

On-site
AED 180,000 - 300,000
DevOps / Infrastructure / SRE / Platform Engineering | Systemsltd | Dubai, Onsite
DevOps / Infrastructure / SRE / Platform Engineering | Systemsltd | Dubai, Onsite

Tech Junction Ltd • Dubai

On-site
AED 350,000 - 600,000
Senior Site Reliability Engineer — FinTech Cloud & DevOps
Senior Site Reliability Engineer — FinTech Cloud & DevOps

Epergne Solutions • Dubai

On-site
AED 200,000 - 300,000
Site Reliability Engineer (SRE) - Azure focus
Site Reliability Engineer (SRE) - Azure focus

Dicetek LLC • Dubai

On-site
AED 300,000 - 550,000
Senior Site Reliability Engineer FinTech Cloud & Kubernetes
Senior Site Reliability Engineer FinTech Cloud & Kubernetes

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000