Site Reliability Engineering Manager

Open Innovation AI

Abu Dhabi Emirate

On-site

AED 150,000 - 210,000

Full time

39 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Open Innovation AI is seeking an experienced SRE Manager to lead the L2 Support - Site Reliability Engineering team overseeing GPU-dense, on-premises AI platforms. You will own SLA performance, drive incident response, and coordinate cross-functional escalation and recovery efforts.

The role requires 8+ years in L2/L3 support or SRE, strong people-management skills, and fluent English. You will report to the Head of Technical Operations and ensure 24/7 readiness and continuous improvement across

Qualifications

  • Bachelor’s degree in computer science, IT, engineering, or a related field.
  • 8+ years of experience in L2/L3 support, SRE, systems engineering, or infrastructure operations, including large-scale on-premises production environments, with at least 3 years of direct people-management responsibility.
  • Proven people-management experience covering hiring, onboarding, objective setting, performance management, coaching, career development, and management of underperformance.
  • Experience managing customer-facing production services in a multi-customer environment and balancing operational priorities, service commitments, risk, and available engineering capacity.
  • Demonstrated major-incident leadership, including technical coordination, recovery governance, stakeholder communication, escalation, and post-incident corrective-action management.
  • Experience owning SLAs, operational KPIs, on-call coverage, backlog governance, service reviews, and continuous-improvement plans.
  • Fluent written and spoken English

Responsibilities

  • Manage the L2 Support - Site Reliability Engineering team day to day, including workload allocation across customers, environments, services, and issues, and rebalance assignments as priorities and operational risks change.
  • Own L2 service performance and maturity, including SLA compliance, operational KPIs, backlog health, ticket ageing, repeat incidents, escalation quality, and customer-specific support readiness.
  • Coordinate the team's incident response so that every incident has the right technical owner, is tracked through recovery and resolution, and is escalated promptly when deeper expertise or additional authority is required.
  • Act as, or appoint, the technical incident lead for P1 and P2 incidents, ensuring clear technical ownership, coordinated recovery, timely stakeholder updates, evidence preservation, and post-incident follow-up.
  • Plan and own on-call rotations and shift coverage to provide sustainable 24/7 continuity for key accounts
  • Provide technical oversight and judgement on complex incidents by understanding the problem, assessing risk and options, and deciding on priority, assignment, recovery approach, and escalation, while relying on senior engineers for deep hands-on troubleshooting.
  • Ensure consistent, ITIL-aligned Incident, Problem, and Change Management across the L2 function, including change risk assessment, execution readiness, rollback planning, and follow-up of corrective actions.

Skills

People management
Incident leadership
SRE practices
Communication skills

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Linux
Kafka
Redis
PostgreSQL
ITIL framework

Job description

Open Innovation AI is a global technology company that specializes in developing advanced solutions for managing AI workloads. Its flagship product, the Open Innovation Cluster Manager (OICM), orchestrates complex AI tasks efficiently across diverse infrastructures. The platform is hardware-agnostic, optimized for various GPUs and accelerators hardware, and facilitates seamless integration and scalability for enterprise AI applications. Open Innovation AI focuses on optimizing and simplifying AI workload management and making AI technologies accessible to organizations of all sizes. With its innovative solutions, companies can reduce operational costs, accelerate time to value, and maximize their return on investment, ensuring that their AI strategies contribute directly to enhanced business outcomes.

Role Overview:

The SRE Manager leads Open Innovation AI's L2 Support - Site Reliability Engineering team, a multidisciplinary function responsible for the reliable operation of GPU-dense, on-premises AI platforms across compute, storage, networking, virtualization, and Kubernetes. The role owns the people, processes, technical coordination, service performance, and continuous improvement required to deliver reliable L2 support across customer environments. This is a management and operational-leadership role. The holder must be technically credible enough to understand complex incidents, assess priority and risk, guide recovery, and make sound assignment and escalation decisions, without needing to be the deepest hands-on specialist in every domain. The role is accountable for SLA performance, 24/7 coverage, ITIL-aligned practices, operational readiness, reliability improvement, and the development of the engineering team, working closely with the Service Desk, Service Delivery Managers, L3, product teams, delivery teams, customers, and technology partners

Role Responsibilities:

  • Manage the L2 Support - Site Reliability Engineering team day to day, including workload allocation across customers, environments, services, and issues, and rebalance assignments as priorities and operational risks change.
  • Own L2 service performance and maturity, including SLA compliance, operational KPIs, backlog health, ticket ageing, repeat incidents, escalation quality, and customer-specific support readiness.
  • Coordinate the team's incident response so that every incident has the right technical owner, is tracked through recovery and resolution, and is escalated promptly when deeper expertise or additional authority is required.
  • Act as, or appoint, the technical incident lead for P1 and P2 incidents, ensuring clear technical ownership, coordinated recovery, timely stakeholder updates, evidence preservation, and post-incident follow-up.
  • Plan and own on-call rotations and shift coverage to provide sustainable 24/7 continuity for key accounts
  • Provide technical oversight and judgement on complex incidents by understanding the problem, assessing risk and options, and deciding on priority, assignment, recovery approach, and escalation, while relying on senior engineers for deep hands-on troubleshooting.
  • Ensure consistent, ITIL-aligned Incident, Problem, and Change Management across the L2 function, including change risk assessment, execution readiness, rollback planning, and follow-up of corrective actions.
  • Support Service Delivery Managers in customer operational reviews, escalations, SLA analysis, service-improvement plans, and communication of technical risks, while maintaining clear boundaries between technical operations and commercial account ownership.
  • Coordinate closely with the Service Desk for smooth ticket flow and accurate triage, and with L3, product, and delivery teams on root-cause analysis, fix validation, knowledge transfer, and operational readiness.
  • Own the L2 operational-readiness assessment for new releases, platforms, and customer environments, including monitoring, access, documentation, runbooks, training, escalation paths, recovery procedures, rollback plans, and known limitations.
  • Oversee operational health through regular review of system status, cluster integrity, capacity, utilization, availability, performance, platform behavior, and emerging risks across supported environments.
  • Coordinate technical escalations with vendors, delivery partners, customer infrastructure teams, and other third parties when incidents or risks cross organizational boundaries.
  • Ensure that the team maintains accurate and usable SOPs, runbooks, troubleshooting guides, knowledge-base articles, architecture references, support matrices, and shift-handover records.
  • Lead all people-management activities for the team, including hiring, onboarding, objective setting, performance management, coaching, succession planning, skills development, and building sufficient technical depth and redundancy.
  • Prepare operational reports and P1/P2 post-incident reports with clear root-cause analysis, customer impact, timeline, corrective actions, owners, and due dates.
  • Keep the Head of Technical Operations and relevant stakeholders informed about service performance, team capacity, operational events, material risks, dependencies, and continuous-improvement initiatives.
  • Exercise the authority required to meet operational commitments, including reprioritizing work, reassigning engineers, escalating resource conflicts, requiring incident reviews, recommending emergency changes, and rejecting operationally unready service handovers

Required experience & Qualification

  • Bachelor’s degree in computer science, Information Technology, Engineering, or a related field.
  • 8 or more years of experience in L2/L3 support, SRE, systems engineering, or infrastructure operations, including large-scale on-premises production environments, with at least 3 years of direct people-management responsibility.
  • Proven people-management experience covering hiring, onboarding, objective setting, performance management, coaching, career development, and management of underperformance.
  • Experience managing customer-facing production services in a multi-customer environment and balancing operational priorities, service commitments, risk, and available engineering capacity.
  • Demonstrated major-incident leadership, including technical coordination, recovery governance, stakeholder communication, escalation, and post-incident corrective-action management.
  • Experience owning SLAs, operational KPIs, on-call coverage, backlog governance, service reviews, and continuous-improvement plans.
  • Proven ability to build, scale, or mature an SRE, infrastructure-operations, or technical-support function, including establishing operating practices, technical ownership, skills coverage, and knowledge management.
  • Technically credible across key SRE domains including GPU-dense compute, Linux, Kubernetes, high-performance networking over Ethernet, InfiniBand and RoCE, distributed storage, and virtualization. The candidate must be able to understand incidents, assess priority and risk, and make informed assignment and escalation decisions; deep hands-on expertise in every layer is not required.
  • Working knowledge of distributed middleware and data-layer technologies such as Kafka, Redis, and PostgreSQL.
  • Solid knowledge of ITIL-aligned Incident, Problem, and Change Management, including operational risk assessment and change-readiness governance.
  • Strong coordination, decision-making, and communication skills, with the ability to maintain a clear view of ownership, priorities, dependencies, and risk and to communicate effectively with engineers, customers, partners, and senior leadership.
  • Experience working in secure, isolated, air-gapped, or compliance-driven on-premises environments.
  • Ability to obtain and maintain the clearance required for regular access to security-controlled customer sites.
  • Fluent written and spoken English

Preferred Skills:

  • Hands-on experience with Kubernetes and HPC or AI platform-management tooling.
  • Experience coordinating technical escalations with hardware, storage, network, platform, or software vendors.
  • Relevant certifications such as ITIL, CKA or CKAD, RHCE or RHCA, CCNP, or VMware VCP.
  • Reporting To: Head of Technical Operations
  • Manages: The L2 Support - Site Reliability Engineering team, spanning compute, storage, networking, virtualization, Kubernetes, platform operations, and reliability engineering.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Manager, On-Prem AI Platform Reliability
SRE Manager, On-Prem AI Platform Reliability

Open Innovation AI • Abu Dhabi Emirate

On-site
AED 150,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Good co India • United Arab Emirates

On-site
AED 240,000 - 480,000
Service Desk Lead - Application Support
Service Desk Lead - Application Support

Open Innovation AI • Abu Dhabi

On-site
AED 268,000 - 469,000
Head of Site Reliability Engineering (SRE)
Head of Site Reliability Engineering (SRE)

Client of Mark Williams • Dubai

On-site
AED 600,000 - 1,200,000
Head of Site Reliability Engineering (SRE)
Head of Site Reliability Engineering (SRE)

Mark Williams Recruitment • Abu Dhabi

On-site
AED 380,000 - 700,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

31 CONCEPT • United Arab Emirates

On-site
AED 300,000 - 460,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Loft Orbital Solutions • Abu Dhabi

On-site
AED 450,000 - 750,000
Site Reliability Engineer - AIOps
Site Reliability Engineer - AIOps

DiceTek UAE • Al Ruways Industrial City

On-site
Vice President Site Reliability Engineering
Vice President Site Reliability Engineering

Remotedxb • Dubai

On-site
AED 480,000 - 720,000
SisaInfosec- Product Support Engineer
SisaInfosec- Product Support Engineer

Nexthire • Sharjah

On-site
AED 279,000 - 446,000