Senior SRE Specialist: Observability & Incident Response

Takamol Holding

Riyadh

On-site

SAR 150,000 - 270,000

Full time

37 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Takamol Holding is seeking an IT Operations/SRE professional to support incidents across digital platforms and to monitor the Elastic Observability stack. You will collaborate with Platform Engineering and development teams to resolve issues within SLAs and contribute to automation initiatives.

Ideal candidates bring 1–3 years in IT ops or SRE, strong scripting skills, and a good grasp of monitoring concepts, ITIL practices, and modern cloud-native environments.

Qualifications

  • Bachelor’s degree in Computer Science, IT, Engineering, or related field (or equivalent experience).
  • 1–3 years of experience in IT operations, system administration, application support, DevOps, or SRE.
  • Familiarity with Observability tools such as Elastic Stack (Elasticsearch, Kibana, etc.), including basic querying and dashboard usage.
  • Knowledge of Linux systems and scripting (Bash, Python, or Go).
  • Understanding of monitoring, logging, and alerting concepts.
  • Experience with ITSM tools (ServiceNow, Jira, Zendesk) and ITIL practices.
  • Strong grasp of incident, problem, and change management.
  • Basic experience with cloud native environments and containers such as Docker and Kubernetes.
  • Strong critical thinking, troubleshooting, and communication skills.

Responsibilities

  • Provide support for application incidents across digital platforms, working with Platform Engineering, App Development, and customer support to resolve per SLAs and escalation procedures.
  • Operate and monitor the Elastic Observability stack — Elasticsearch, Kibana, Fleet Server, APM Server, Elastic Agent via ECK on OKE.
  • Assist with day-to-day Elasticsearch operations like ILM, SLM, data tier housekeeping, and capacity monitoring.
  • Troubleshoot telemetry ingestion issues across logs, metrics, traces, and synthetic monitors.
  • Maintain and update Kibana dashboards, alerting rules, and saved objects under SRE guidance.
  • Perform root cause analysis and participate in blameless post-incident reviews to improve reliability.
  • Collaborate with Platform Engineering to automate tasks, improve pipelines, and enhance observability using Terraform, Helm charts, and scripting.
  • Develop and maintain support documentation, runbooks, and knowledge base articles aligned to incident response procedures.
  • Manage incidents and requests via Jira/ServiceNow, ensuring documentation in the service management system.
  • Participate in on-call rotation and reduce operational toil through automation and tooling.
  • Monitor KPIs like MTTD and MTTR and report on incident metrics.
  • Collaborate with cross-functional teams and vendors to improve reliability and security posture.

Skills

Observability tools
Linux scripting
Monitoring & alerting
ITSM familiarity
Docker & Kubernetes
Incident management
Communication skills
Problem solving

Education

Bachelor’s degree in Computer Science/IT/Engineering

Tools

Elastic Stack
Jira
ServiceNow
Zendesk
Terraform
Helm
Bash
Python
Go
Docker
Kubernetes

Job description

Takamol Holding is seeking an IT Operations/SRE professional to support incidents across digital platforms and to monitor the Elastic Observability stack. You will collaborate with Platform Engineering and development teams to resolve issues within SLAs and contribute to automation initiatives.

Ideal candidates bring 1–3 years in IT ops or SRE, strong scripting skills, and a good grasp of monitoring concepts, ITIL practices, and modern cloud-native environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer Specialist
Senior Site Reliability Engineer Specialist

Takamol Holding • Riyadh

On-site
SAR 150,000 - 270,000
Expert Site Reliability Engineer
Expert Site Reliability Engineer

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
Senior DevOps & SRE Engineer - Cloud & Reliability
Senior DevOps & SRE Engineer - Cloud & Reliability

Norconsult Telematics • Riyadh

On-site
SAR 350,000 - 520,000
Senior SRE: Remote, Scale High-Throughput Systems
Senior SRE: Remote, Scale High-Throughput Systems

Jobgether SRL • Saudi Arabia

On-site
SAR 300,000 - 600,000
Fully remote
Global distributed team
Ownership over production reliability
+2
Senior Elasticsearch Observability Engineer
Senior Elasticsearch Observability Engineer

Emdad By Elm • Jeddah

On-site
SAR 360,000 - 480,000
Senior Manager - Application Operations
Senior Manager - Application Operations

Rasan • Riyadh

On-site
SAR 300,000 - 480,000
Senior DevOps & SRE Engineer: Cloud, Automation, Reliability
Senior DevOps & SRE Engineer: Cloud, Automation, Reliability

Norconsult Telematics Limited • Saudi Arabia

On-site
SAR 280,000 - 420,000
Senior Site Reliability Engineer: Scale, Automation, Observability
Senior Site Reliability Engineer: Scale, Automation, Observability

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
Senior Elasticsearch Observability Architect
Senior Elasticsearch Observability Architect

Emdad By Elm • Jeddah

On-site
SAR 360,000 - 480,000
Senior Manager, Application Operations & Reliability
Senior Manager, Application Operations & Reliability

Rasan • Riyadh

On-site
SAR 300,000 - 480,000