Senior Cloud & Infrastructure Engineer — AI-Driven Ops

Core42

Abu Dhabi

On-site

AED 320,000 - 520,000

Full time

44 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Yearly Bonus
Exclusive Discount Cards: Esaad and FZ
Premium Family Insurance

Job summary

Core42 is seeking a Senior Engineer – Infrastructure & Cloud Engineering to design, deploy, and optimize large-scale private cloud, virtualization, and observability platforms across our infrastructure. The role requires hands-on expertise in OpenStack, OpenShift, and related technologies, plus AI-assisted operations for automated incident detection and remediation.

You will collaborate with architecture, product, SRE, security, and operations teams, supporting capacity planning, upgrades,

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, Software Engineering, or a related technology discipline; or equivalent practical experience.
  • 5+ years of hands-on experience designing, implementing, operating, troubleshooting, and managing private cloud, virtualization, infrastructure, observability, or platform engineering environments.
  • Strong hands-on experience with at least one major cloud or virtualization platform, such as OpenStack, Proxmox, Red Hat OpenShift, or equivalent technologies.
  • Hands-on experience with AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational efficiency, incident response, and service reliability.
  • Strong hands-on experience with compute technologies, including x86 hardware, KVM, Linux operating systems, hypervisors, firmware, server lifecycle management, and orchestration services.
  • Expert-level Linux administration skills, including troubleshooting, performance analysis, system tuning, patching, and operational support of Linux-based infrastructure environments.
  • Strong understanding of hardware architecture and components, including x86/ARM, NUMA, memory channels, NICs, GPU/AI accelerators, firmware, and large-scale server platforms.
  • Good understanding of data center networking concepts, including OSI model, TCP/IP, routing, firewalls, load balancing, VLAN/VXLAN, DNS, DHCP, and related technologies.
  • Hands-on experience with observability concepts and platforms, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, OpenTelemetry, Prometheus, Grafana, ELK/OpenSearch, Splunk, Zabbix, or similar technologies.
  • Experience managing large-scale public or private cloud environments, cloud service provider platforms, mission-critical infrastructure, or high-availability managed service environments is highly desirable.
  • Strong experience with automation, Infrastructure as Code, CI/CD, GitOps, and scripting using Ansible, Terraform, Helm, Jenkins, GitLab CI/CD, Python, Go, Bash, or similar technologies.
  • Understanding of security monitoring, SIEM concepts, compliance requirements, and data governance considerations in cloud and infrastructure environments is an advantage.
  • Relevant certifications in Linux, virtualization, cloud computing, OpenStack, Kubernetes, OpenShift, or ITSM are advantageous.
  • Strong analytical, troubleshooting, communication, documentation, stakeholder management, and problem-solving skills.

Responsibilities

  • Design, implement, and operate observability platforms and services, including metrics, logs, traces, dashboards, alerting, and service health visibility using technologies such as Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, and related platforms.
  • Contribute to the design, implementation, and operation of large-scale private cloud, virtualization, and container platforms based on OpenStack, Red Hat OpenShift, and related infrastructure technologies.
  • Develop, integrate, and maintain AI-powered operational agents that leverage observability, monitoring, logging, ITSM, and platform telemetry systems to automate incident detection, root cause analysis, operational workflows, and approved remediation activities in accordance with established operational processes and governance controls.
  • Collaborate with architecture, product, platform engineering, SRE, security, and operations teams on technology evaluation, integration, solution design, and operational readiness.
  • Support capacity planning, performance optimization, production upgrades, migrations, incident response, root cause analysis, and continuous improvement of platform reliability and operational resilience.
  • Integrate observability and platform management capabilities with ITSM, incident, change, and problem management processes and tools such as Jira, ServiceNow, PagerDuty, Opsgenie, or similar platforms.
  • Collaborate with security teams to ensure compute, virtualization, cloud, observability, and management platforms are secure, hardened, and aligned with cybersecurity, compliance, and data governance requirements.
  • Create and maintain technical documentation, including implementation guides, operational procedures, dashboards, runbooks, diagrams, standards, and knowledge base articles.
  • Prepare and deliver technical knowledge transfer sessions for operational teams, SREs, and engineering stakeholders.
  • Participate in on-call rotations and provide technical escalation support for critical production incidents, major service disruptions, and platform emergencies, ensuring timely restoration of services and effective root cause resolution.
  • Work with process and operations teams to improve support workflows, service onboarding, operational procedures, and collaboration efficiency.

Skills

Analytical skills
Troubleshooting
Communication
Documentation
Stakeholder management
Problem solving

Education

Bachelor’s or Master’s degree

Tools

Prometheus
Grafana
OpenTelemetry
ELK/OpenSearch
Splunk
Zabbix
Jira
ServiceNow
PagerDuty

Job description

Core42 is seeking a Senior Engineer – Infrastructure & Cloud Engineering to design, deploy, and optimize large-scale private cloud, virtualization, and observability platforms across our infrastructure. The role requires hands-on expertise in OpenStack, OpenShift, and related technologies, plus AI-assisted operations for automated incident detection and remediation.

You will collaborate with architecture, product, SRE, security, and operations teams, supporting capacity planning, upgrades,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Container Platform Engineer — Kubernetes & OpenShift
Senior Container Platform Engineer — Kubernetes & OpenShift

Core42 • United Arab Emirates

On-site
AED 300,000 - 420,000
Competitive salary
Yearly bonus
Discount cards (Esaad/Fazaa)
+2
Senior HPC Operations Engineer for AI/ML Workloads
Senior HPC Operations Engineer for AI/ML Workloads

Core42 • Abu Dhabi

On-site
AED 350,000 - 650,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
+2
Senior HPC Engineer: Build & Optimize AI Clusters
Senior HPC Engineer: Build & Optimize AI Clusters

Core42 • Abu Dhabi

On-site
AED 420,000 - 650,000
Competitive Salary
Yearly Bonus
Discount Cards Esaad and Fazaa
+2
Senior DevOps Engineer — Cloud, CI/CD & Kubernetes Leader
Senior DevOps Engineer — Cloud, CI/CD & Kubernetes Leader

Core42 • United Arab Emirates

On-site
AED 350,000 - 600,000
Competitive Salary
Yearly Bonus
Discount Cards
+2
Senior Java Backend Engineer - Cloud & OpenStack Expert
Senior Java Backend Engineer - Cloud & OpenStack Expert

Core42 • Dubai

On-site
AED 350,000 - 500,000
Yearly bonus
Discount cards (Esaad/Fazaa)
Premium family insurance
+1
Lead Engineer – Infrastructure & Cloud Engineering
Lead Engineer – Infrastructure & Cloud Engineering

SmartChoice International GCC • Abu Dhabi

On-site
AED 335,000 - 580,000
Senior DevOps Engineer: Cloud, CI/CD & Automation Leader
Senior DevOps Engineer: Cloud, CI/CD & Automation Leader

Core42 • Abu Dhabi

On-site
AED 300,000 - 520,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards
+2
Senior Engineer - Infrastructure and Cloud Engineering
Senior Engineer - Infrastructure and Cloud Engineering

Core42 • Abu Dhabi

On-site
AED 320,000 - 520,000
Yearly Bonus
Exclusive Discount Cards: Esaad and FZ
Premium Family Insurance
Senior Backend Engineer (Go) — Cloud-Native AI Infra
Senior Backend Engineer (Go) — Cloud-Native AI Infra

Core42 • Abu Dhabi

On-site
AED 320,000 - 520,000
Competitive Salary
Yearly Bonus
Discount Cards (Esaad & Fazaa)
+2
Senior Engineer - Container Platforms
Senior Engineer - Container Platforms

Core42 • United Arab Emirates

On-site
AED 300,000 - 420,000
Competitive salary
Yearly bonus
Discount cards (Esaad/Fazaa)
+2