SRE Lead

Dizzaract

Abu Dhabi

On-site

AED 420,000 - 680,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
24 days annual leave
Office in Yas Creative Hub
Continuous learning

Job summary

Dizzaract is seeking an SRE Lead to take end-to-end ownership of infrastructure reliability, observability, automation, and production engineering for FAR Labs. You will lead architecture, incident response, and technical standards across backend and platform teams.

In this hands-on leadership role, you will champion reliability, scalability, and cost efficiency while collaborating with security, product, and engineering leadership to keep systems resilient as the business grows in Abu Dhabi and

Qualifications

  • 7+ years of experience in SRE, DevOps, or platform engineering.

Responsibilities

  • Own reliability, availability, and performance of FAR Labs production infrastructure.

Skills

Kubernetes mastery
Terraform IaC
CI/CD pipelines
Incident management
Observability tooling
Python/Go scripting
Linux networking
Cloud-native security

Tools

Kubernetes
Terraform
Ansible
Prometheus
Grafana
ELK/EFK

Job description

01. ABOUT THE COMPANY

Dizzaract is a product-driven company operating at the intersection of gaming, digital platforms, and AI. We build and scale multiple products, including FAR Labs & Gamed — each exploring a different space, yet united by a shared approach: moving fast, staying curious, and focusing on things that people actually use.

We operate as a collaborative, non-hierarchical team where ideas are valued based on their impact, not their origin, and where AI is embedded across everything we build, from infrastructure to product decisions.

02. ABOUT THE ROLE

FAR Labs is building a distributed inference platform designed to serve large language models efficiently across diverse hardware. Serving quality — latency, throughput, memory efficiency, cost, and model quality — sits at the core of the product.

As our systems grow in complexity and scale, reliability becomes an engineering discipline of its own. We are looking for an SRE Lead who can take end-to-end ownership of infrastructure reliability, observability, automation, and production engineering across FAR Labs.

This is a hands-on technical leadership role. You will own the reliability, scalability, and operational maturity of our production environment while defining the SRE standards used across engineering.

You will work closely with Backend, Platform, Product, Security, and Engineering leadership to ensure our systems remain reliable as the business scales.

We are looking for someone who can move comfortably between architecture, hands-on engineering, incident response, automation, and technical leadership.

03. WHAT YOU'LL OWN
Infrastructure & Reliability
  • Own the reliability, availability, scalability, and performance of FAR Labs production infrastructure.
  • Design and evolve resilient infrastructure supporting growing and distributed workloads.
  • Identify infrastructure bottlenecks, single points of failure, and operational risks.
  • Establish capacity planning, backup, recovery, and disaster-recovery practices.
  • Drive improvements in infrastructure performance, stability, and cost efficiency.
Kubernetes & Cloud Infrastructure
  • Own the architecture, operation, and optimization of Kubernetes environments.
  • Improve cluster reliability, networking, resource utilization, deployment strategies, and workload management.
  • Build and maintain infrastructure through Terraform, Helm, Ansible, or equivalent IaC tooling.
  • Ensure infrastructure is automated, reproducible, version-controlled, and scalable.
  • Work deeply with Linux, Docker/containerd, networking, storage, and cloud-native infrastructure.
CI/CD Automation
  • Design and improve production-grade CI/CD pipelines.
  • Automate infrastructure provisioning, deployments, configuration, testing, and operational workflows.
  • Improve deployment safety through automated validation, rollback mechanisms, and appropriate release controls.
  • Reduce manual operational work through automation.
Observability
  • Own the observability strategy across FAR Labs.
  • Build monitoring, logging, tracing, dashboards, and alerting infrastructure.
  • Work with Prometheus, Grafana, VictoriaMetrics, ELK/EFK, or equivalent technologies.
  • Establish visibility into infrastructure health, application health, performance, and failures.
  • Improve alert quality and reduce unnecessary operational noise.
SRE Standarts
  • Define and implement SLIs, SLOs, SLAs, and error budgets.
  • Establish reliability standards for critical services.
  • Partner with engineering teams to make reliability part of the development lifecycle.
  • Use operational data to identify recurring problems and prioritize improvements.
Incident Management
  • Lead technical response to critical production incidents.
  • Establish incident management and escalation processes.
  • Lead RCA and blameless post-mortems.
  • Translate incidents into concrete engineering improvements.
  • Participate in and help structure the on-call rotation.
Security & Operational Readiness
  • Embed infrastructure and cloud security best practices.
  • Strengthen access management, secrets management, network security, CI/CD security, and runtime security.
  • Support security, compliance, and audit-readiness requirements where relevant.
Technical Leadership
  • Set the technical direction for SRE and infrastructure engineering.
  • Establish engineering standards and best practices.
  • Review infrastructure architecture and major technical changes.
  • Mentor engineers and increase infrastructure expertise across teams.
  • Remain hands-on with implementation and troubleshooting.
04. REQUIREMENTS
  • 7+ years of experience in SRE, DevOps, Platform Engineering, Infrastructure Engineering, or similar roles.
  • Strong experience owning production infrastructure at scale.
  • Deep hands-on expertise with Kubernetes.
  • Strong experience with Terraform and Infrastructure as Code.
  • Experience designing and operating production-grade CI/CD pipelines.
  • Strong knowledge of Linux, networking, DNS, load balancing, storage, containers, and distributed systems.
  • Experience with Prometheus, Grafana, VictoriaMetrics, ELK/EFK, or similar observability stacks.
  • Strong scripting/programming skills in Python, Go, Bash, or similar.
  • Practical experience with SLIs, SLOs, SLAs, alerting, and error budgets.
  • Experience leading production incidents, RCA, and post-mortems.
  • Strong understanding of cloud-native and infrastructure security.
  • Experience making architectural decisions and driving technical standards.
  • Strong ownership, communication, and problem-solving skills.
05. NICE TO HAVE
  • Experience supporting AI/ML or HPC workloads.
  • Experience with GPU infrastructure, scheduling, resource management, or performance optimization.
  • Experience operating large-scale distributed systems.
  • Multi-cloud, hybrid-cloud, or self-managed infrastructure experience.
  • Experience optimizing infrastructure costs and capacity.
  • Experience with service mesh, distributed tracing, or advanced Kubernetes networking.
  • Previous experience as an SRE Lead, Platform Lead, Infrastructure Lead, or Tech Lead.
06. WHAT WE OFFER
  • Real technical ownership of site reliability, infrastructure, and production operations across FAR Labs.
  • The opportunity to define the SRE strategy, standards, and technical direction as our platform scales.
  • Hands-on work with technically challenging problems across Kubernetes, distributed systems, cloud infrastructure, networking, observability, and automation.
  • The opportunity to build and evolve high-availability, scalable production infrastructure rather than simply maintain existing systems.
  • Direct influence over architecture, SLOs, incident management, CI/CD, observability, and infrastructure performance.
  • A small, senior engineering environment where you can set technical standards, mentor engineers, and see your decisions translate directly into platform and product outcomes.
  • Fast execution, low bureaucracy, and a highly collaborative, idea-driven team.
  • Competitive salary with performance-based incentives.
  • 24 days annual leave, plus public holidays.
  • Health insurance.
  • Modern office in Yas Creative Hub.
  • Continuous learning through real-world problem solving — not just theory.
  • The opportunity to influence both technical architecture and product direction.
  • A diverse, open-minded team where ideas are genuinely heard.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Backend Engineer
Backend Engineer

Dizzaract • Abu Dhabi

On-site
AED 300,000 - 420,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
Senior Engineering Manager
Senior Engineering Manager

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 660,000
SRE Lead: Platform Reliability & Infra Mastery
SRE Lead: Platform Reliability & Infra Mastery

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 680,000
Health insurance
24 days annual leave
Office in Yas Creative Hub
+1
Lead AI Engineer (Inference Serving & Performance)
Lead AI Engineer (Inference Serving & Performance)

Dizzaract • Abu Dhabi

On-site
AED 420,000 - 720,000
Health insurance
24 days annual leave
Modern office in Yas Creative Hub
Senior Engineering Manager at Dizzaract FZ LLC
Senior Engineering Manager at Dizzaract FZ LLC

Dizzaract FZ LLC • Abu Dhabi

On-site
AED 279,000 - 335,000
24 days annual leave
Health insurance
Modern office in Yas Creative Hub
Associate ML Ops Engineer
Associate ML Ops Engineer

Tanqeeb • Abu Dhabi

On-site
AED 120,000 - 240,000
Health insurance
Visa sponsorship
21 days annual leave
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether SRL • United Arab Emirates

On-site
AED 350,000 - 600,000
Fully remote
Global engineering org
Ownership over reliability
+2
Senior DevOps / Site Reliability Engineer (SRE)
Senior DevOps / Site Reliability Engineer (SRE)

Stellar Technologies • Abu Dhabi

On-site
AED 360,000 - 540,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NEXT Ventures • Dubai

On-site
AED 240,000 - 360,000
SRE (Site Reliability Engineer)
SRE (Site Reliability Engineer)

Dicetek LLC • Abu Dhabi

On-site
AED 180,000 - 300,000