Senior Site Reliability Engineer Specialist

Takamol Holding

Riyadh

On-site

SAR 150,000 - 270,000

Full time

38 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Takamol Holding is seeking an IT Operations/SRE professional to support incidents across digital platforms and to monitor the Elastic Observability stack. You will collaborate with Platform Engineering and development teams to resolve issues within SLAs and contribute to automation initiatives.

Ideal candidates bring 1–3 years in IT ops or SRE, strong scripting skills, and a good grasp of monitoring concepts, ITIL practices, and modern cloud-native environments.

Qualifications

  • Bachelor’s degree in Computer Science, IT, Engineering, or related field (or equivalent experience).
  • 1–3 years of experience in IT operations, system administration, application support, DevOps, or SRE.
  • Familiarity with Observability tools such as Elastic Stack (Elasticsearch, Kibana, etc.), including basic querying and dashboard usage.
  • Knowledge of Linux systems and scripting (Bash, Python, or Go).
  • Understanding of monitoring, logging, and alerting concepts.
  • Experience with ITSM tools (ServiceNow, Jira, Zendesk) and ITIL practices.
  • Strong grasp of incident, problem, and change management.
  • Basic experience with cloud native environments and containers such as Docker and Kubernetes.
  • Strong critical thinking, troubleshooting, and communication skills.

Responsibilities

  • Provide support for application incidents across digital platforms, working with Platform Engineering, App Development, and customer support to resolve per SLAs and escalation procedures.
  • Operate and monitor the Elastic Observability stack — Elasticsearch, Kibana, Fleet Server, APM Server, Elastic Agent via ECK on OKE.
  • Assist with day-to-day Elasticsearch operations like ILM, SLM, data tier housekeeping, and capacity monitoring.
  • Troubleshoot telemetry ingestion issues across logs, metrics, traces, and synthetic monitors.
  • Maintain and update Kibana dashboards, alerting rules, and saved objects under SRE guidance.
  • Perform root cause analysis and participate in blameless post-incident reviews to improve reliability.
  • Collaborate with Platform Engineering to automate tasks, improve pipelines, and enhance observability using Terraform, Helm charts, and scripting.
  • Develop and maintain support documentation, runbooks, and knowledge base articles aligned to incident response procedures.
  • Manage incidents and requests via Jira/ServiceNow, ensuring documentation in the service management system.
  • Participate in on-call rotation and reduce operational toil through automation and tooling.
  • Monitor KPIs like MTTD and MTTR and report on incident metrics.
  • Collaborate with cross-functional teams and vendors to improve reliability and security posture.

Skills

Observability tools
Linux scripting
Monitoring & alerting
ITSM familiarity
Docker & Kubernetes
Incident management
Communication skills
Problem solving

Education

Bachelor’s degree in Computer Science/IT/Engineering

Tools

Elastic Stack
Jira
ServiceNow
Zendesk
Terraform
Helm
Bash
Python
Go
Docker
Kubernetes

Job description

Job Description
  • Provide support for application incidents across digital platforms, working closely with Platform Engineering, Application Development, and customer support teams to ensure timely resolution according to established SLAs and escalation procedures.
  • Operate and monitor the Elastic Observability stack — including Elasticsearch cluster health, Kibana, Fleet Server, APM Server, and Elastic Agent — deployed and managed via ECK on OKE.
  • Assist with day-to-day Elasticsearch operations such as index lifecycle management (ILM), snapshot lifecycle management (SLM), data tier housekeeping (hot, warm, cold, frozen), and capacity monitoring.
  • Troubleshoot telemetry ingestion issues across logs, metrics, traces, and synthetic monitors, ensuring consistent data collection from all platforms.
  • Maintain and update Kibana dashboards, alerting rules, and saved objects under the guidance of the SRE Manager.
  • Perform root cause analysis and participate in blameless post-incident reviews to improve system reliability and reduce recurrence.
  • Collaborate with Platform Engineering to automate repetitive tasks, improve deployment pipelines, and enhance observability coverage using Terraform, Helm charts, and scripting.
  • Develop and maintain support documentation, runbooks, and knowledge base articles aligned to standardized incident response procedures.
  • Manage and prioritize incidents and requests via the ticketing system (Jira/ServiceNow), ensuring all incidents, requests, and resolutions are documented in the service management system.
  • Participate in an on-call rotation and help reduce operational toil through automation and tooling.
  • Monitor and report on key performance metrics related to incident management, including mean time to detect (MTTD) and mean time to resolve (MTTR).
  • Collaborate with cross-functional teams and vendor partners to improve overall system reliability, observability maturity, and security posture.
Job Requirements
  • Bachelor’s degree in Computer Science, IT, Engineering, or related field (or equivalent experience).
  • 1–3 years of experience in IT operations, system administration, application support, DevOps, or SRE.
  • Familiarity with Observbility tools such as Elastic Stack (Elasticsearch, Kibana, etc.), including basic querying and dashboard usage.
  • Knowledge of Linux systems and scripting (Bash, Python, or Go).
  • Understanding of monitoring, logging, and alerting concepts.
  • Experience with ITSM tools (ServiceNow, Jira, Zendesk) and ITIL practices.
  • Strong grasp of incident, problem, and change management.
  • Basic experience with cloud native enviroments and containers such as Docker and Kubernetes.
  • Strong critical thinking, troubleshooting, and communication skills.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Expert Site Reliability Engineer
Expert Site Reliability Engineer

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
Senior SRE Specialist: Observability & Incident Response
Senior SRE Specialist: Observability & Incident Response

Takamol Holding • Riyadh

On-site
SAR 150,000 - 270,000
Site Reliability Engineer
Site Reliability Engineer

Lucidya | لوسيديا • Riyadh

On-site
SAR 250,000 - 360,000
Senior Manager - Application Operations
Senior Manager - Application Operations

Rasan • Riyadh

On-site
SAR 300,000 - 480,000
Senior DevOps Engineer
Senior DevOps Engineer

Norconsult Telematics • Riyadh

On-site
SAR 350,000 - 520,000
Senior Elasticsearch Observability Engineer
Senior Elasticsearch Observability Engineer

Emdad By Elm • Jeddah

On-site
SAR 360,000 - 480,000
Senior DevOps Engineer
Senior DevOps Engineer

Norconsult Telematics Limited • Saudi Arabia

On-site
SAR 280,000 - 420,000
Senior DevOps Engineer
Senior DevOps Engineer

Norconsult Telematics • Medina

On-site
SAR 600,000 - 900,000
Expert Platform Engineer
Expert Platform Engineer

TAWANTECH • Riyadh

On-site
SAR 180,000 - 300,000
Automation & Observability Engineer
Automation & Observability Engineer

HCLTech • Riyadh

On-site
SAR 180,000 - 280,000