SRE Manager, On-Prem AI Platform Reliability

Open Innovation AI

Abu Dhabi Emirate

On-site

AED 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Open Innovation AI is seeking an experienced SRE Manager to lead the L2 Support - Site Reliability Engineering team overseeing GPU-dense, on-premises AI platforms. You will own SLA performance, drive incident response, and coordinate cross-functional escalation and recovery efforts.

The role requires 8+ years in L2/L3 support or SRE, strong people-management skills, and fluent English. You will report to the Head of Technical Operations and ensure 24/7 readiness and continuous improvement across

Qualifications

  • Bachelor’s degree in computer science, IT, engineering, or a related field.
  • 8+ years of experience in L2/L3 support, SRE, systems engineering, or infrastructure operations, including large-scale on-premises production environments, with at least 3 years of direct people-management responsibility.
  • Proven people-management experience covering hiring, onboarding, objective setting, performance management, coaching, career development, and management of underperformance.
  • Experience managing customer-facing production services in a multi-customer environment and balancing operational priorities, service commitments, risk, and available engineering capacity.
  • Demonstrated major-incident leadership, including technical coordination, recovery governance, stakeholder communication, escalation, and post-incident corrective-action management.
  • Experience owning SLAs, operational KPIs, on-call coverage, backlog governance, service reviews, and continuous-improvement plans.
  • Fluent written and spoken English

Responsibilities

  • Manage the L2 Support - Site Reliability Engineering team day to day, including workload allocation across customers, environments, services, and issues, and rebalance assignments as priorities and operational risks change.
  • Own L2 service performance and maturity, including SLA compliance, operational KPIs, backlog health, ticket ageing, repeat incidents, escalation quality, and customer-specific support readiness.
  • Coordinate the team's incident response so that every incident has the right technical owner, is tracked through recovery and resolution, and is escalated promptly when deeper expertise or additional authority is required.
  • Act as, or appoint, the technical incident lead for P1 and P2 incidents, ensuring clear technical ownership, coordinated recovery, timely stakeholder updates, evidence preservation, and post-incident follow-up.
  • Plan and own on-call rotations and shift coverage to provide sustainable 24/7 continuity for key accounts
  • Provide technical oversight and judgement on complex incidents by understanding the problem, assessing risk and options, and deciding on priority, assignment, recovery approach, and escalation, while relying on senior engineers for deep hands-on troubleshooting.
  • Ensure consistent, ITIL-aligned Incident, Problem, and Change Management across the L2 function, including change risk assessment, execution readiness, rollback planning, and follow-up of corrective actions.

Skills

People management
Incident leadership
SRE practices
Communication skills

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Linux
Kafka
Redis
PostgreSQL
ITIL framework

Job description

Open Innovation AI is seeking an experienced SRE Manager to lead the L2 Support - Site Reliability Engineering team overseeing GPU-dense, on-premises AI platforms. You will own SLA performance, drive incident response, and coordinate cross-functional escalation and recovery efforts.

The role requires 8+ years in L2/L3 support or SRE, strong people-management skills, and fluent English. You will report to the Head of Technical Operations and ensure 24/7 readiness and continuous improvement across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Manager
Site Reliability Engineering Manager

Open Innovation AI • Abu Dhabi Emirate

On-site
AED 150,000 - 210,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Dicetek LLC • Abu Dhabi

Hybrid
AED 120,000 - 180,000
Head of Site Reliability Engineering (SRE)
Head of Site Reliability Engineering (SRE)

Client of Mark Williams • Dubai

On-site
AED 600,000 - 1,200,000
Site Reliability Engineer (SRE) - Azure AI
Site Reliability Engineer (SRE) - Azure AI

Dicetek LLC • Abu Dhabi

On-site
AED 260,000 - 420,000
AI Platform SRE: Resilience & Automation
AI Platform SRE: Resilience & Automation

Dicetek LLC • Abu Dhabi

Hybrid
AED 120,000 - 180,000
AI Platform SRE for Banking-Grade Resilience
AI Platform SRE for Banking-Grade Resilience

Deeplight • Abu Dhabi

On-site
AED 260,000 - 460,000
Visa sponsorship
Comprehensive health insurance
Professional development support
+3
On-Site SRE Lead: Reliability & Growth in Dubai
On-Site SRE Lead: Reliability & Growth in Dubai

HCLTech • United Arab Emirates

On-site
AED 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Dubai

On-site
AED 200,000 - 300,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000
Site Reliability Engineer (SRE) - Azure focus
Site Reliability Engineer (SRE) - Azure focus

Dicetek LLC • Dubai

On-site
AED 300,000 - 550,000