Site Reliability Engineer

Open Innovation AI

Abu Dhabi

On-site

AED 240,000 - 400,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Open Innovation AI seeks a Site Reliability Engineer to support deployments across customer environments, including secure on‑prem infrastructure. You will troubleshoot hardware, Linux, Kubernetes, and middleware while maintaining service availability and program integrity.

You will work within Incident, Change, and Problem Management processes, document procedures, and collaborate with L1/L3 teams to resolve issues and optimize platform performance.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 4–7 years of experience in SRE, DevOps, Infrastructure Operations, or Platform Engineering in on-prem or secure environments.
  • Strong proficiency in Linux system administration, troubleshooting, log analysis, service management, and performance tuning.
  • Hands-on experience with Kubernetes, container runtimes, and distributed on‑prem systems.
  • Understanding of compute, storage, networking, and virtualization in enterprise installs.
  • Experience with middleware/data-layer components (Kafka, Redis, PostgreSQL) in distributed on‑prem setups.
  • Knowledge of ITIL-aligned operations and structured processes.
  • Ability to diagnose complex issues across multiple stack layers.
  • Experience in secure/restricted environments is an advantage.
  • Excellent communication and documentation skills.
  • Certifications such as RHCSA/RHCE, CKA/CKAD/CKS.

Responsibilities

  • Maintain and support deployments in customer environments including on-prem and air-gapped setups.
  • Ensure availability, upgrades, and stability of AI platforms and solutions.
  • Diagnose incidents across hardware, Linux, Kubernetes, containers, and middleware.
  • Analyze logs and system behavior to determine root causes and restore service.
  • Follow Change Management procedures for system updates and upgrades.
  • Understand OICM architecture and customer usage patterns.
  • Collaborate with L1/Service Desk for issue triage and guidance.
  • Escalate complex issues to L3 with detailed analysis.
  • Perform on-site health checks of Kubernetes clusters and resources.
  • Work with Systems Engineering to resolve performance issues across compute, storage, networking, and Kubernetes layers.
  • Update SOPs, runbooks, and known-issues documentation.
  • Contribute to post-incident reviews and improvement actions.
  • Adhere to Incident, Change, and Problem Management processes.

Skills

Linux administration
Incident management
Problem management
Analytical thinking
Communication
Troubleshooting
Documentation

Education

Bachelor's degree in Computer Science / IT / Engineering

Tools

Kubernetes
Container runtimes
Kafka
Redis
PostgreSQL
Virtualization

Job description

Company Overview

Open Innovation AI is a global technology company that specializes in developing advanced solutions for managing AI workloads. Its flagship product, the Open Innovation Cluster Manager (OICM), orchestrates complex AI tasks efficiently across diverse infrastructures. The platform is hardware-agnostic, optimized for various GPUs and accelerators hardware, and facilitates seamless integration and scalability for enterprise AI applications. Open Innovation AI focuses on optimizing and simplifying AI workload management and making AI technologies accessible to organizations of all sizes. With its innovative solutions, companies can reduce operational costs, accelerate time to value, and maximize their return on investment, ensuring that their AI strategies contribute directly to enhanced business outcomes.

Role Overview

The Site Reliability Engineer is responsible for supporting and maintaining Open Innovation AI Products and deployments across customer environments, including secure and isolated on-premises infrastructures. This role requires strong troubleshooting skills across hardware, Linux OS, Kubernetes, middleware, and application layers.

The engineer is expected to diagnose and resolve technical incidents, applying deep product knowledge and strong analytical skills to restore service availability. The role requires solid understanding of operational processes such as Incident, Change, and Problem Management, along with a thorough grasp of the product architecture and how customers use it in production environments.

Role Responsibilities
  • Experienced in Customer production environment deployments, mostly restricted connectivity and air gapped
  • Involves ensuring availability, upgrades, stability of end-to-end AI platforms and solutions
  • Diagnose and resolve incidents across hardware, Linux OS, Kubernetes clusters, containerized services, middleware, and platform components.
  • Perform detailed analysis of logs, system behavior, and application output to identify root causes and restore service functionality.
  • Review, validate, and execute approved changes following Change Management procedures, including system updates, configuration adjustments, and component upgrades.
  • Maintain a strong understanding of the OICM and other OI product's architecture, its services, dependencies, and typical customer usage patterns.
  • Collaborate with L1 and Service Desk teams by providing technical guidance, clarifying issue details, and ensuring accurate ticket triage.
  • Escalate complex, code-level or product-defect issues to L3 with complete diagnostic information and structured analysis.
  • Conduct on-site platform health assessments, validating Kubernetes cluster status, service integrity, system resources, and overall environment readiness.
  • Work closely with the Systems Engineering team to analyze and resolve performance issues across compute, storage, networking, and Kubernetes layers, and ensure that identified optimizations are reflected in the product and operational practices.
  • Update and maintain technical documentation including SOPs, runbooks, troubleshooting steps, and known-issue guides.
  • Participate in post-incident reviews, contributing technical insights and recommending improvements to prevent recurrence.
  • Ensure all activities adhere to established Incident, Change, and Problem Management processes.
Required experience & Qualification
  • Bachelor's degree in computer science, Information Technology, Engineering, or a related field.
  • 4–7 years of experience in SRE, DevOps, Infrastructure Operations, or Platform Engineering roles within on-prem or secure environments.
  • Strong proficiency in Linux system administration, including troubleshooting, log analysis, service management, and performance tuning.
  • Hands-on experience with Kubernetes, container runtimes, and distributed systems deployed in on-prem environments.
  • Solid understanding of compute, storage, networking, and virtualization layers relevant to enterprise installations.
  • Practical experience with middleware and data-layer components such as Kafka, Redis, PostgreSQL, or similar technologies used in distributed on-prem environments.
  • Strong understanding of ITIL-aligned and experience operating within structured operational frameworks.
  • Ability to diagnose complex issues across multiple layers of the stack.
  • Experience working in secure, restricted, or isolated environments is an advantage.
  • Excellent analytical skills, communication abilities, and a methodical approach to troubleshooting.
  • Ability to produce clear technical documentation, including SOPs, runbooks, and investigation reports.
  • Certifications such as RHCSA/RHCE, CKA/CKAD/CKS.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Delivery Engineer
Senior Delivery Engineer

Open Innovation AI • Abu Dhabi

On-site
AED 360,000 - 600,000
Forward Deployed Engineer
Forward Deployed Engineer

Open Innovation AI • Abu Dhabi

On-site
AED 360,000 - 600,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Dicetek LLC • Dubai

On-site
AED 480,000 - 780,000
On-Prem SRE: AI Platform Reliability & Kubernetes
On-Prem SRE: AI Platform Reliability & Kubernetes

Open Innovation AI • Abu Dhabi

On-site
AED 240,000 - 400,000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Client of Salt • Abu Dhabi

On-site
AED 320,000 - 520,000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Salt • Abu Dhabi

On-site
AED 480,000 - 720,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Dicetek LLC • Abu Dhabi

On-site
AED 180,000 - 280,000
Site Reliability Engineer (SRE) - Azure focus
Site Reliability Engineer (SRE) - Azure focus

Dicetek LLC • Dubai

On-site
AED 300,000 - 550,000
Senior Engineer - Site Reliability
Senior Engineer - Site Reliability

AIQ • Abu Dhabi

On-site
AED 300,000 - 420,000
Healthcare
Education support for dependents
Leave benefits
Senior DevOps / SRE Engineer
Senior DevOps / SRE Engineer

Legend Holding Group Ltd • Dubai

On-site
AED 450,000 - 650,000