Senior Site Reliability Engineer — Obs, SLOs & Resilience

NEXT Ventures

Dubai

On-site

AED 240,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NEXT Ventures in Dubai is seeking a Senior Site Reliability Engineer to own centralized logging, alerting, and reliability across the platform. You will work in the Platform Engineering squad to keep services observable, fast, and resilient, with a focus on reducing incident duration and noise.

The role requires 5–7 years engineering, 3+ years SRE/DevOps, strong experience with ELK/OpenSearch, Datadog, Kubernetes (EKS preferred), Terraform, and scripting in Python/Bash/Go.

Qualifications

  • 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or reliability‑focused DevOps.
  • Strong hands‑on experience with centralized log management platforms — ingestion, parsing, logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
  • Ability to diagnose incidents through log analysis across Linux and Windows environments, isolating root causes under pressure.
  • Designs alerting systems with tiered thresholds to minimize noise while catching real problems early, with clear escalation paths.
  • Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
  • Experienced defining and operating SLOs, SLIs, and error budgets for customer‑facing services.
  • Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
  • Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
  • Hands‑on experience operating services on Kubernetes; EKS preferred.
  • Familiar with Infrastructure‑as‑Code tooling such as Terraform.
  • Evidence‑driven — profiles and measures before guessing; validates every optimization against before/after data.
  • Reliability‑oriented — treats detection speed and signal quality as first‑class engineering problems.
  • Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.

Responsibilities

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
  • Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.
  • Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers.
  • Define, implement, and own SLOs, SLIs, and error budgets for critical customer‑facing services, and drive improvements against them.
  • Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
  • Lead reliability and performance root‑cause analysis and drive durable fixes.
  • Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.
  • Participate in on‑call rotation with blameless post‑incident reviews, reducing manual toil through automation.

Skills

SRE
Observability
Incident response
Python/Bash/Go
Datadog
Kubernetes
Terraform
ELK/OpenSearch
AI tooling
Linux/Windows

Tools

Datadog Logs
ELK/OpenSearch
Kubernetes (EKS)
Terraform

Job description

NEXT Ventures in Dubai is seeking a Senior Site Reliability Engineer to own centralized logging, alerting, and reliability across the platform. You will work in the Platform Engineering squad to keep services observable, fast, and resilient, with a focus on reducing incident duration and noise.

The role requires 5–7 years engineering, 3+ years SRE/DevOps, strong experience with ELK/OpenSearch, Datadog, Kubernetes (EKS preferred), Terraform, and scripting in Python/Bash/Go.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Engineer: Reliability, Observability & Cloud Automation
SRE Engineer: Reliability, Observability & Cloud Automation

Dicetek LLC • Dubai

On-site
AED 300,000 - 520,000
Site Reliability Engineer SRE
Site Reliability Engineer SRE

Dicetek LLC • Dubai

On-site
AED 300,000 - 520,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NEXT Ventures • Dubai

On-site
AED 240,000 - 360,000
Site Reliability Engineer SRE
Site Reliability Engineer SRE

D4 Insight • Abu Dhabi

On-site
AED 180,000 - 320,000
Staff Software Engineer — AI Reliability & SRE Lead
Staff Software Engineer — AI Reliability & SRE Lead

Remotedxb • Dubai

On-site
AED 350,000 - 750,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Good co India • United Arab Emirates

On-site
AED 240,000 - 480,000
Senior SRE - AIOps & Cloud Reliability Engineer
Senior SRE - AIOps & Cloud Reliability Engineer

DiceTek UAE • Al Ruways Industrial City

On-site
Head of Site Reliability Engineering (SRE)
Head of Site Reliability Engineering (SRE)

Client of Mark Williams • Dubai

On-site
AED 600,000 - 1,200,000
Site Reliability Engineer - Digital Banking
Site Reliability Engineer - Digital Banking

DiceTek UAE • Al Ruways Industrial City

On-site
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Dubai

On-site
AED 200,000 - 300,000