Site Reliability Engineer

Lucidya | لوسيديا

Riyadh

On-site

SAR 250,000 - 360,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lucidya is seeking a Site Reliability Engineer to own cloud infrastructure reliability and scalable deployments. You will design highly available systems, manage workloads across AWS/GCP/Azure, and drive automation across the stack.

You will operate Kubernetes in production, refine monitoring, and work with DevOps to continuously improve performance and resilience. The role emphasizes proactive problem solving and scalable engineering practices.

Qualifications

  • 3+ years in SRE/DevOps or infrastructure engineering, experience at scale.
  • Hands-on with cloud environments (AWS, GCP, or Azure).
  • Production Kubernetes experience and incident troubleshooting.
  • Experience building observability and alerting pipelines.

Responsibilities

  • Design, build, and maintain highly available, scalable infrastructure.
  • Manage workloads across cloud providers and optimize costs.
  • Operate Kubernetes clusters in production with reliable deployments.
  • Implement monitoring using Prometheus, Grafana, Datadog, or ELK.
  • Automate repetitive tasks with scripts and IaC.
  • Collaborate with DevOps and engineering teams to improve performance.

Skills

Kubernetes
Terraform
Docker
Python
CI/CD
Networking
Monitoring

Tools

Terraform
Docker
Kubernetes
GitHub Actions
Jenkins

Job description

About Lucidya

Lucidya is an AI-native platform for customer experience (CX) intelligence that manages entire customer lifecycles autonomously, from initial engagement through retention and growth.

Lucidya is an AI-native platform for customer experience (CX) intelligence that manages entire customer lifecycles autonomously, from initial engagement through retention and growth. Unlike platforms that only surface insights and leave the action to you, Lucidya closes the loop with proprietary NLU technology built in-house and trained on millions of multilingual conversations. This enables marketing, support, CX, and research teams to deliver personalized experiences that drive measurable improvements in customer satisfaction, retention and lifetime value. As we continue scaling globally, the reliability, performance, and resilience of our infrastructure become mission-critical to everything we do.

Why this role matters

At Lucidya, our platform processes massive volumes of real-time customer data. Any downtime, latency, or instability directly impacts our customers' ability to make decisions and serve their own users. This role exists to make sure that doesn't happen. As a Site Reliability Engineer, you'll sit at the heart of our platform's stability, owning the reliability of our cloud infrastructure and ensuring it scales seamlessly as we grow. You won't just react to issues; you'll anticipate them, design systems that prevent them, and build automation that removes them entirely. If you enjoy solving complex infrastructure challenges, eliminating inefficiencies, and building systems that "just work" - this is where you'll thrive.

What You'll Do
You'll make reliability the default
  • You'll design and maintain infrastructure that is highly available, fault-tolerant, and scalable
  • You'll proactively identify and eliminate single points of failure before they become incidents
  • You'll ensure our production systems remain stable, even under increasing scale and load
You'll own and optimize our cloud environments
  • You'll manage and continuously improve workloads across AWS, GCP, or Azure
  • You'll use Infrastructure as Code (Terraform) to standardize and scale infrastructure
  • You'll optimize resource usage to balance performance and cost
You'll run and improve Kubernetes in production
  • You'll operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence
  • You'll troubleshoot issues quickly and ensure smooth deployments and upgrades
  • You'll ensure our containerized workloads perform reliably at scale
You'll build strong observability and respond to incidents
  • You'll implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK
  • You'll define alerting that is meaningful, not noisy
  • You'll respond to incidents, lead root cause analysis, and ensure we learn from every failure
You'll automate everything that shouldn't be manual
  • You'll write scripts and build tooling to eliminate repetitive operational work
  • You'll continuously improve infrastructure efficiency through automation
  • You'll promote a culture where manual work is a temporary state, not the norm
You'll collaborate to improve the entire system
  • You'll work closely with DevOps and engineering teams to solve performance bottlenecks
  • You'll contribute to CI/CD improvements and deployment reliability
  • You'll help shape reliability best practices across the organization
What success looks like (First 90 Days)
First 30 days:
  • You've built a strong understanding of our infrastructure, systems, and workflows
  • You're contributing to day-to-day operations with support from the team
  • You've started identifying areas for improvement in automation and reliability
By 90 days:
  • You're independently managing infrastructure tasks and troubleshooting issues
  • You're actively contributing to reliability and scalability improvements
  • You've taken ownership of parts of our infrastructure and are improving them
Requirements
This is what will make you successful in this role
  • You've spent :3 years working in SRE, DevOps, or infrastructure engineering, and you've seen what breaks at scale
  • You're comfortable working in cloud environments like AWS, GCP, or Azure—and you understand how distributed systems behave
  • You've worked hands-on with Kubernetes in production and know how to troubleshoot it when things go wrong
  • You don't just fix issues - you ask why they happened and make sure they don't happen again
Technically, you likely
  • Use Terraform (or similar IaC tools) to manage infrastructure
  • Work confidently with Docker and Kubernetes
  • Write scripts in Python, Bash, or similar to automate workflows
  • Understand CI/CD pipelines (Jenkins, GitHub Actions, Bitbucket, etc.)
  • Have a solid grasp of networking, load balancing, and high-availability design
When it comes to monitoring
  • You've implemented tools like Prometheus, Grafana, Datadog, or ELK
  • You know the difference between useful alerts and noise
  • You focus on signals that actually drive action
What sets you apart
  • You take ownership - you don't wait to be told something is broken
  • You're calm under pressure and methodical during incidents
  • Simplify complexity instead of adding to it
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer: Scale, Automate, Stabilize Cloud
Site Reliability Engineer: Scale, Automate, Stabilize Cloud

Lucidya | لوسيديا • Riyadh

On-site
SAR 250,000 - 360,000
Expert Site Reliability Engineer
Expert Site Reliability Engineer

TAWANTECH • Riyadh

On-site
SAR 240,000 - 320,000
10x Software Engineer
10x Software Engineer

Lucidya | لوسيديا • Riyadh

On-site
SAR 180,000 - 250,000
10x Software Engineer
10x Software Engineer

Lucidya • Riyad Al Khabra

On-site
SAR 250,000 - 480,000
10x Software Engineer
10x Software Engineer

Lucidya • Riyadh

On-site
SAR 224,971 - 299,962
Senior Site Reliability Engineer Specialist
Senior Site Reliability Engineer Specialist

Takamol Holding • Riyadh

On-site
SAR 150,000 - 270,000
Frontend Software Engineer - Saudi Only
Frontend Software Engineer - Saudi Only

Lucidya | لوسيديا • Jeddah

On-site
SAR 60,000 - 90,000
Frontend Software Engineer - Saudi Only
Frontend Software Engineer - Saudi Only

Lucidya • Jeddah

On-site
SAR 60,000 - 100,000
DevOps Engineer
DevOps Engineer

azmtalent • Saudi Arabia

On-site
SAR 200,000 - 350,000
Senior DevOps Engineer
Senior DevOps Engineer

Norconsult Telematics • Riyadh

On-site
SAR 350,000 - 520,000