Senior Site Reliability Engineer — Obs, SLOs & Resilience

NEXT Ventures

Dubai

On-site

AED 240,000 - 360,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NEXT Ventures in Dubai is seeking a Senior Site Reliability Engineer to own centralized logging, alerting, and reliability across the platform. You will work in the Platform Engineering squad to keep services observable, fast, and resilient, with a focus on reducing incident duration and noise.

The role requires 5–7 years engineering, 3+ years SRE/DevOps, strong experience with ELK/OpenSearch, Datadog, Kubernetes (EKS preferred), Terraform, and scripting in Python/Bash/Go.

Qualifications

  • 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or reliability‑focused DevOps.
  • Strong hands‑on experience with centralized log management platforms — ingestion, parsing, logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
  • Ability to diagnose incidents through log analysis across Linux and Windows environments, isolating root causes under pressure.
  • Designs alerting systems with tiered thresholds to minimize noise while catching real problems early, with clear escalation paths.
  • Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
  • Experienced defining and operating SLOs, SLIs, and error budgets for customer‑facing services.
  • Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
  • Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
  • Hands‑on experience operating services on Kubernetes; EKS preferred.
  • Familiar with Infrastructure‑as‑Code tooling such as Terraform.
  • Evidence‑driven — profiles and measures before guessing; validates every optimization against before/after data.
  • Reliability‑oriented — treats detection speed and signal quality as first‑class engineering problems.
  • Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.

Responsibilities

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
  • Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.
  • Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers.
  • Define, implement, and own SLOs, SLIs, and error budgets for critical customer‑facing services, and drive improvements against them.
  • Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
  • Lead reliability and performance root‑cause analysis and drive durable fixes.
  • Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.
  • Participate in on‑call rotation with blameless post‑incident reviews, reducing manual toil through automation.

Skills

SRE
Observability
Incident response
Python/Bash/Go
Datadog
Kubernetes
Terraform
ELK/OpenSearch
AI tooling
Linux/Windows

Tools

Datadog Logs
ELK/OpenSearch
Kubernetes (EKS)
Terraform

Job description

NEXT Ventures in Dubai is seeking a Senior Site Reliability Engineer to own centralized logging, alerting, and reliability across the platform. You will work in the Platform Engineering squad to keep services observable, fast, and resilient, with a focus on reducing incident duration and noise.

The role requires 5–7 years engineering, 3+ years SRE/DevOps, strong experience with ELK/OpenSearch, Datadog, Kubernetes (EKS preferred), Terraform, and scripting in Python/Bash/Go.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Lead Avrioc Technologies On-site Fast Track available
Site Reliability Engineer - Lead Avrioc Technologies On-site Fast Track available

HireHouse • Abu Dhabi

On-site
AED 300,000 - 600,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NEXT Ventures • Dubai

On-site
AED 240,000 - 360,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Dicetek LLC • Dubai

On-site
AED 480,000 - 780,000
SRE: Observability, Automation & Incident Response (Abu Dhabi)
SRE: Observability, Automation & Incident Response (Abu Dhabi)

Innovations Global • Abu Dhabi

On-site
AED 350,000 - 550,000
Site Reliability Engineer - Automation & Observability
Site Reliability Engineer - Automation & Observability

Dicetek LLC • Dubai

On-site
AED 300,000 - 550,000
Site Reliability Engineer (SRE) - Azure focus
Site Reliability Engineer (SRE) - Azure focus

Dicetek LLC • Dubai

On-site
AED 300,000 - 550,000
Observability Engineer
Observability Engineer

Reqiva • Dubai

Hybrid
AED 279,000 - 446,000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Client of Salt • Abu Dhabi

On-site
AED 320,000 - 520,000
SRE (Site Reliability Engineer)
SRE (Site Reliability Engineer)

Dicetek LLC • Abu Dhabi

On-site
AED 180,000 - 300,000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Salt • Abu Dhabi

On-site
AED 480,000 - 720,000