Sr. Site Reliability Engineer (SRE)

sifiapp

Riyadh

On-site

SAR 260,000 - 420,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SiFi is a rapidly growing B2B Fin-Tech company in Saudi Arabia, seeking a Senior Site Reliability Engineer to own reliability, performance, and scalability of production systems. You will design, automate, and operate mission-critical environments including Kubernetes clusters, database DR, and cross-region networking.

The role emphasizes automation, proactive incident management, and collaboration with development teams to embed resilience in releases.

Qualifications

  • 5+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering.
  • Deep experience with networking, load balancing, and DNS.
  • Experience integrating alerting and monitoring systems with collaboration tools (e.g., Teams or Slack).
  • OCI, Kubernetes (OKE), Terraform, Python & PowerShell are required.

Responsibilities

  • Maintain multi-region cloud infrastructure using Terraform-based IaC.
  • Operate and optimize Kubernetes (OKE) clusters running microservices and data pipelines.
  • Manage SQL Server backup/restore pipelines, DR testing, and performance tuning.
  • Ensure high availability for .NET and Python applications behind load balancers and WAF.

Tools

OCI
Kubernetes (OKE)
Microsoft SQL Server
Terraform
Python
PowerShell

Job description

Job Description
About SiFi

SiFi is a rapidly growing B2B Fin-Tech company transforming expense management for businesses in Saudi Arabia. As a licensed EMI from the Saudi Central Bank, we empower companies with innovative tools to simplify finance management.

Position Overview

We are looking for a Senior Site Reliability Engineer (SRE) who will take ownership of the reliability, performance, and scalability of our production systems. You will design, automate, and operate mission-critical environments that include Kubernetes clusters, database disaster recovery, workflow orchestration, and multi-region networking.

This role suits engineers who think deeply about systems — combining infrastructure, automation, and diagnostic reasoning to drive operational excellence.

Primary Responsibilities
Reliability, Availability & Infrastructure
  • Maintain and evolve multi-region cloud infrastructure using Terraform-based Infrastructure as Code (IaC).
  • Operate and optimize Kubernetes (OKE) clusters running microservices, data pipelines, and workflow orchestration.
  • Manage SQL Server backup/restore pipelines, DR testing, and performance optimization.
  • Ensure high availability for .NET and Python applications hosted behind load balancers and WAF.
  • Design and maintain cross-network connectivity (DRGs, LPGs, VCNs, subnets, and NSGs).
Observability & Automation
  • Build and maintain a centralized orchestration platform integrated with alerting and notification systems.
  • Develop self-healing, monitoring, and auto-remediation scripts for infrastructure and databases.
  • Implement logging, metrics, and tracing pipelines.
  • Automate recurring operational tasks using Python, Bash, and PowerShell to reduce manual effort and improve reliability.
DevOps, CI/CD & Security
  • Manage GitHub Actions and Octopus Deploy pipelines for backend and data services.
  • Apply strong security principles — least privilege, network segmentation, secure credentials, and encrypted communications.
  • Promote GitOps and Infrastructure-as-Code practices to ensure repeatable and traceable deployments.
  • Collaborate with developers to embed reliability and resilience into every release.
Collaboration & Incident Management
  • Lead incident response, run blameless post-mortems, and turn findings into lasting improvements.
  • Partner closely with engineering teams to drive design and code-level reliability improvements.
  • Conduct capacity planning, cost optimization, and system tuning for performance and scalability.
  • Mentor engineers in automation, observability, and root-cause analysis best practices.
Troubleshooting Mindset & Diagnostic Thinking

We value engineers who:

  • Approach issues systematically and validate assumptions with data.
  • Treat incidents as opportunities to improve design and automation.
  • Rely on metrics, logs, and tracing rather than guesswork.
  • Communicate findings clear&nobreak;ly and document learnings for future reference.
  • Continuously refine how problems are detected, escalated and resolved.
Requirements
  • 5+ years of experience in Site Reliability, DevOps, or Infrastructure Engineering.

Deep experience with:

  • Solid understanding of networking, load balancing, and DNS.
  • Proven ability to analyze incidents and automate resolution.
  • Experience integrating alerting and monitoring systems with communication tools (e.g., Microsoft Teams or Slack).
  • Oracle Cloud Infrastructure (OCI) (compute, networking, storage, monitoring)
  • Kubernetes (OKE) — deployments, ingress controllers, autoscaling
  • Microsoft SQL Server — backup/restore automation, DR planning, performance tuning
  • Terraform — multi-region and cross-tenant infrastructure automation
  • Python & PowerShell — automation and system scripting
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer (SRE)
Sr. Site Reliability Engineer (SRE)

SiFi - Simplified Financial Solutions Company • Riyadh

On-site
SAR 350,000 - 550,000
Senior SRE: Scalable Cloud, Automation & Reliability
Senior SRE: Scalable Cloud, Automation & Reliability

SiFi - Simplified Financial Solutions Company • Riyadh

On-site
SAR 350,000 - 550,000
Senior SRE: Cloud, Kubernetes & Automation Leader
Senior SRE: Cloud, Kubernetes & Automation Leader

sifiapp • Riyadh

On-site
SAR 260,000 - 420,000
Expert Site Reliability Engineer
Expert Site Reliability Engineer

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
Senior Site Reliability Engineer Specialist
Senior Site Reliability Engineer Specialist

Takamol Holding • Riyadh

On-site
SAR 150,000 - 270,000
Senior DevOps Engineer
Senior DevOps Engineer

Norconsult Telematics Limited • Saudi Arabia

On-site
SAR 280,000 - 420,000
Senior Site Reliability Engineer: Scale, Automation, Observability
Senior Site Reliability Engineer: Scale, Automation, Observability

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
Site Reliability Engineer
Site Reliability Engineer

Lucidya | لوسيديا • Riyadh

On-site
SAR 250,000 - 360,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobgether • Saudi Arabia

Remote
SAR 180,000 - 300,000
Fully remote in Saudi Arabia
Competitive compensation
Technical leadership and mentoring
+1
Senior SRE: Remote, Scale High-Throughput Systems
Senior SRE: Remote, Scale High-Throughput Systems

Jobgether SRL • Saudi Arabia

On-site
SAR 300,000 - 600,000
Fully remote
Global distributed team
Ownership over production reliability
+2