Expert Site Reliability Engineer

TAWANTECH

Riyadh

On-site

SAR 240,000 - 320,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TAWANTECH, based in Riyadh, seeks an experienced Site Reliability Engineer to drive reliability across critical technology services, using automation, observability, and cloud platforms. You will define SLIs/SLOs, lead incident response, and mentor teams in SRE practices.

Strong expertise in Kubernetes, Terraform, and CI/CD is required, with a focus on high availability and capacity planning within production environments. On-site in Saudi Arabia with opportunities to influence architecture.

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, IT, or related field.
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
  • Strong experience in cloud platforms, Kubernetes, and production environments.
  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
  • Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
  • Proven experience in complex incident management, troubleshooting, and RCA.
  • Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
  • Experience driving reliability improvements and reducing operational toil through automation.
  • Strong analytical, problem-solving, and technical leadership skills.
  • Experience in Banking, FinTech, or Payment environments is preferred.

Responsibilities

  • Define and implement advanced reliability engineering practices across critical technology services.
  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
  • Design automation to reduce manual operational activities and improve system resilience.
  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
  • Lead technical analysis and resolution of complex production incidents.
  • Conduct root-cause analysis and drive permanent corrective and preventive actions.
  • Design solutions to improve system availability, scalability, capacity, and disaster resilience.
  • Identify reliability risks and recommend architectural and engineering improvements.
  • Drive performance engineering and capacity planning for critical services.
  • Provide advanced technical guidance and mentorship on SRE practices.
  • Promote automation and engineering approaches that reduce operational toil and improve service reliability.

Skills

Cloud platforms
Kubernetes
Monitoring & observability
SRE practices
Automation & scripting
CI/CD & IaC
Incident management
Root Cause Analysis
Capacity planning
Technical leadership
Banking/FinTech domain

Education

Bachelor's degree in Computer Science / Software Engineering / IT

Tools

Kubernetes
Terraform
CI/CD

Job description

Purpose:

To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.

Main Duties and Responsibilities:

  • Define and implement advanced reliability engineering practices across critical technology services.
  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
  • Design automation to reduce manual operational activities and improve system resilience.
  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
  • Lead technical analysis and resolution of complex production incidents.
  • Conduct root-cause analysis and drive permanent corrective and preventive actions.
  • Design solutions to improve system availability, scalability, capacity, and disaster resilience.
  • Identify reliability risks and recommend architectural and engineering improvements.
  • Drive performance engineering and capacity planning for critical services.
  • Provide advanced technical guidance and mentorship on SRE practices.
  • Promote automation and engineering approaches that reduce operational toil and improve service reliability.
QUALIFICATIONS & REQUIREMENTS
  • Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles.
  • Strong experience in cloud platforms, Kubernetes, and production environments.
  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
  • Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
  • Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
  • Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
  • Experience driving reliability improvements and reducing operational toil through automation.
  • Strong analytical, problem-solving, and technical leadership skills.
  • Experience in Banking, FinTech, or Payment environments is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

PrimeGate for Communications and IT • Riyadh

On-site
SAR 224,000 - 300,000
Expert Platform Engineer
Expert Platform Engineer

TAWANTECH • Riyadh

On-site
SAR 180,000 - 300,000
Senior Site Reliability Engineer: Scale, Automation, Observability
Senior Site Reliability Engineer: Scale, Automation, Observability

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
Tech Lead
Tech Lead

TAWANTECH • Riyadh

On-site
SAR 240,000 - 420,000
Senior Developer
Senior Developer

TAWANTECH • Riyadh

On-site
SAR 180,000 - 280,000
Expert Backend Engineer
Expert Backend Engineer

TAWANTECH • Riyadh

On-site
SAR 300,000 - 420,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Deepsource Technologies • Riyadh

On-site
SAR 201,000 - 312,000
Reliability Engineer
Reliability Engineer

KBR, Inc. • Al Jubayl

On-site
Application Reliability & Incident Engineer
Application Reliability & Incident Engineer

DS DeepSource • Riyadh

On-site
SAR 167,000 - 279,000
Manager - Technical Infrastructure Specialist
Manager - Technical Infrastructure Specialist

Datamatics Technologies • Eastern Province

On-site
SAR 350,000 - 700,000