Senior Site Reliability Engineer

NEXT Ventures

Dubai

On-site

AED 240,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NEXT Ventures in Dubai is seeking a Senior Site Reliability Engineer to own centralized logging, alerting, and reliability across the platform. You will work in the Platform Engineering squad to keep services observable, fast, and resilient, with a focus on reducing incident duration and noise.

The role requires 5–7 years engineering, 3+ years SRE/DevOps, strong experience with ELK/OpenSearch, Datadog, Kubernetes (EKS preferred), Terraform, and scripting in Python/Bash/Go.

Qualifications

  • 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or reliability‑focused DevOps.
  • Strong hands‑on experience with centralized log management platforms — ingestion, parsing, logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
  • Ability to diagnose incidents through log analysis across Linux and Windows environments, isolating root causes under pressure.
  • Designs alerting systems with tiered thresholds to minimize noise while catching real problems early, with clear escalation paths.
  • Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
  • Experienced defining and operating SLOs, SLIs, and error budgets for customer‑facing services.
  • Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
  • Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
  • Hands‑on experience operating services on Kubernetes; EKS preferred.
  • Familiar with Infrastructure‑as‑Code tooling such as Terraform.
  • Evidence‑driven — profiles and measures before guessing; validates every optimization against before/after data.
  • Reliability‑oriented — treats detection speed and signal quality as first‑class engineering problems.
  • Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.

Responsibilities

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
  • Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.
  • Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers.
  • Define, implement, and own SLOs, SLIs, and error budgets for critical customer‑facing services, and drive improvements against them.
  • Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
  • Lead reliability and performance root‑cause analysis and drive durable fixes.
  • Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.
  • Participate in on‑call rotation with blameless post‑incident reviews, reducing manual toil through automation.

Skills

SRE
Observability
Incident response
Python/Bash/Go
Datadog
Kubernetes
Terraform
ELK/OpenSearch
AI tooling
Linux/Windows

Tools

Datadog Logs
ELK/OpenSearch
Kubernetes (EKS)
Terraform

Job description

Who We Are

NEXT Ventures is a global fintech group powering FundedNext — one of the world's fastest-growing proprietary trading platforms — and FNmarkets, a regulated CFD brokerage. Across offices in Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai, we build and operate the technology that lets traders access global markets at scale. Our Platform Engineering team owns the infrastructure, reliability, and observability backbone that every product squad depends on.

Your Role in Our Mission

As our Site Reliability Engineer, you are the dedicated specialist who keeps our services observable, fast, and resilient. You own the centralized logging and alerting backbone, drive service‑level optimization across the stack, and perform log analysis across both Linux and Windows environments. Working within the Platform Engineering squad, you execute the reliability initiatives that free the squad lead to focus on architecture — and you are the reason incidents are short, signals are clean, and detection is fast.

This role is distinct from our DevSecOps Engineer: DevSecOps builds and secures the platform foundation; you measure, detect, and optimize on top of it.

How You'll Make An Impact
Centralized Logging & Alerting
  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
  • Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.
Service‑Level Optimization & Reliability
  • Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers — profile before guessing, measure every fix.
  • Define, implement, and own SLOs, SLIs, and error budgets for critical customer‑facing services, and drive improvements against them.
  • Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
  • Lead reliability and performance root‑cause analysis and drive durable fixes.
  • Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.
Operations & Cross‑Team Support
  • Participate in a shared on‑call rotation with solid runbooks and blameless post‑incident reviews.
  • Continuously reduce manual toil through automation and better tooling.
  • Collaborate with product squads to instrument services, define meaningful SLIs, and surface the right signals.
  • Document runbooks, dashboards, and operational procedures so any engineer can respond to incidents with clear guidance.
What You Bring
  • 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability‑focused DevOps work.
  • Strong hands‑on experience with centralized log management platforms — ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
  • Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure.
  • Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths.
  • Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
  • Experienced defining and operating SLOs, SLIs, and error budgets for customer‑facing services.
  • Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
  • Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
  • Hands‑on experience operating services on Kubernetes; EKS preferred.
  • Familiar with Infrastructure‑as‑Code tooling such as Terraform for collaboration with the DevSecOps team.
  • Evidence‑driven — profiles and measures before guessing; validates every optimization against before/after data.
  • Reliability‑oriented — treats detection speed and signal quality as first‑class engineering problems.
  • Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.
X‑Factor: AI‑Native Engineering
  • You actively use modern AI agentic workflows daily — not limited to Copilot autocomplete.
  • You are proficient with Claude Code, Cursor, Windsurf, or equivalent tools.
  • You are comfortable with project‑level AI configuration (CLAUDE.md, rules files), agentic task delegation, and AI‑driven code review.
  • You think in terms of 5–10× productivity through AI‑augmented development — and you can demonstrate it.
Your Journey After Applying
  • Stage 1 – TA Interview
  • Stage 2 – Screening Questionnaire
  • Stage 3 – Hiring Manager Interview
  • Stage 4 – Head of IT Interview
Why Join NEXT
  • Work on reliability challenges at real scale — 100M+ row data stores, multi‑region traffic, high‑frequency trading infrastructure.
  • A team that treats observability and detection speed as first‑class engineering problems, not afterthoughts.
  • Flat structure — your work directly shapes how the platform operates, not filtered through layers of process.
  • Offices across Bangladesh, Malaysia, Sri Lanka, Cyprus, and Dubai — a genuinely global engineering team.
  • Competitive compensation benchmarked to your market, with room to grow as the team scales.

Experience level: Senior

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Database Reliability Engineer
Database Reliability Engineer

NEXT Ventures • Dubai

On-site
AED 360,000 - 480,000
DevSecOps Engineer
DevSecOps Engineer

NEXT Ventures • Dubai

On-site
AED 310,000 - 460,000
Senior Site Reliability Engineer — Obs, SLOs & Resilience
Senior Site Reliability Engineer — Obs, SLOs & Resilience

NEXT Ventures • Dubai

On-site
AED 240,000 - 360,000
Site Reliability Engineer
Site Reliability Engineer

MultiBank Group • Dubai

On-site
AED 280,000 - 420,000
Competitive salary
Site Reliability Engineer - AIOps
Site Reliability Engineer - AIOps

DiceTek UAE • Al Ruways Industrial City

On-site
Site Reliability Engineer - Digital Banking
Site Reliability Engineer - Digital Banking

DiceTek UAE • Al Ruways Industrial City

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Good co India • United Arab Emirates

On-site
AED 240,000 - 480,000
Senior Site Reliability Engineer (Performance and Scalability)
Senior Site Reliability Engineer (Performance and Scalability)

JobCubby • United Arab Emirates

On-site
AED 300,000 - 540,000
Immediate impact on a high-growth team
Top-tier compensation packages
Work with regional talent from top UAE
Trading and Risk Advisor — CFD
Trading and Risk Advisor — CFD

NEXT Ventures • Dubai

On-site
AED 367,000 - 551,000
Fullstack Lead (Laravel & Node.js)
Fullstack Lead (Laravel & Node.js)

NEXT Ventures • Dubai

On-site
AED 160,000 - 230,000