Senior Engineer - Site Reliability

MATRIXJV Co. Ltd.

Abu Dhabi

On-site

AED 300,000 - 540,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Healthcare
Education support
Leave benefits

Job summary

AIQ in Abu Dhabi seeks a Senior Site Reliability Engineer to enhance platform reliability, lead key projects, and improve observability. You will respond to complex production incidents and drive automation across deployment pipelines.

You will mentor junior engineers while aligning with QHSE and governance policies, with a focus on scalable operations and continuous improvement. The role demands strong Linux/Unix, Kubernetes, and cloud experience.

Qualifications

  • Bachelor’s degree in CS, Analytics, or engineering.
  • Masters preferred.
  • +5 years in an SRE/DevOps/Platform role.
  • +5 years managing Kubernetes clusters.
  • +5 years with monitoring/observability platforms.
  • Familiarity with at least one database.

Responsibilities

  • Maintain monitoring, alerting, and incident response systems.
  • Improve observability and fix performance bottlenecks.
  • Contribute to infrastructure automation and pipelines.
  • Drive SLO/SLI adoption with engineering teams.
  • Lead root cause analysis and preventive solutions.
  • Mentor junior engineers and scale operational excellence.

Skills

Linux/Unix
Kubernetes
CI/CD Pipelines
Observability
Python scripting
Bash scripting
Incident response
Cloud platforms
OpenShift

Education

Bachelor’s Degree in Computer Science
Masters Degree preferred

Tools

Docker
Kubernetes
Prometheus
Grafana
ELK
Sentry
Terraform
Ansible
Redis
RabbitMQ
Kafka
OpenShift
Azure
Huawei Cloud

Job description

Job Description:

About the Company:

AIQ is an Abu Dhabi-based technology company that develops and deploys industrial artificial intelligence (AI) technologies at scale, focused on the energy sector. As a venture between Presight (G42) and ADNOC, AIQ has productized 15 AI-enabled solutions that support clients to perform better, protect teams and equipment, keep operations sustainable, and rapidly scale successes. The organization embraces an innovative and entrepreneurial spirit, looking to push boundaries and solve transformational challenges across industry. It welcomes professionals that share the desire to make meaningful and impactful contributions to the mission to deliver responsible AI to the heart of industrial processes. Always on the forefront of technology, AIQ provides its talent with an environment to thrive and excel. Working at AIQ includes participating in some of the most significant industrial transformation projects, interacting with massive data pools, utilizing sophisticated AI infrastructure that is powered by a NVIDIA GPU cloud computing platform, and access to abundant computing, storage, and network resources made available from across the G42 ecosystem.

Overview:

Job Title: Senior Engineer - Site Reliability

Job Location: Abu Dhabi, UAE

About AIQ

AIQ is a joint venture company in the United Arab Emirates, majority-owned by Presight (an ADX-listed G42 company) alongside ADNOC and G42, which focuses on developing artificial intelligence technologies. AIQ develops and commercializes AI products and applications for the oil and gas industry. It aims in providing end-to-end solutions by using its data, cloud and talents to develop AI solutions that seek to reduce costs and generate revenue for its clients.

AIQ embodies an innovative and entrepreneurial spirit that embraces challenges to push boundaries and seeks to welcome professionals to its team that share to desire to make meaningful and impactful contributions to its mission. Always on the cutting edge of technology, AIQ provides its talent all the opportunities to thrive and excel. Working at AIQ includes dealing with massive data sets, an AI infrastructure that is powered by the latest NVIDIA GPU cloud computing platform and access to limitless computing, storage and network resources.

Responsibilities:

Key Responsibilities

As a Senior SRE, you will enhance the reliability and performance of our platforms. You will lead key reliability projects, improve observability, and respond to complex production incidents.

Key Responsibilities:

  • Maintain and evolve monitoring, alerting, and incident response systems.
  • Proactively find and fix performance bottlenecks and failure points.
  • Contribute to infrastructure automation and deployment pipelines.
  • Drive SLO/SLI adoption in collaboration with engineering teams.
  • Lead root cause analysis and build preventive solutions.
  • Mentor junior engineers and help scale operational excellence.
  • Analyze service performance, identify bottlenecks, and provide measurable improvement plans.
  • Maintain the environment’s health by continuously monitoring technical and business metrics, configuring alerts for potential issues, and proactively addressing risks to prevent disruption.
  • Deploy application updates with minimal disruption to services
  • Identify, evaluate, and conduct proof-of-concepts for new technologies.
  • Contribute to the knowledge base.
  • Review and modify CI/CD principles and service maturity iteratively, striving for continuous improvement
  • Comply with QHSE (Quality Health Safety and Environment), Business Continuity, Information Security, Privacy, Risk, Compliance Management, and Governance of Organizations policies, procedures, plans, and related risk assessments.
Qualifications:

Requirements:

Qualifications:

  • Bachelor’s Degree in Business Analytics, Data Science, Computer Science, Engineering, or a related field.
  • Masters Degree is preferred.

Experience:

  • +5 years in a SRE/DevOps/Sysadmin/Platform Engineer role
  • +5 years of experience in managing Kubernetes clusters.
  • +5 years of experience in configuring and using monitoring/observability platforms
  • Familiarity with at least one type of database

Skills

  • Solid experience with containerized environments (Docker, Kubernetes).
  • Hands-on with CI/CD pipelines and automation tools.
  • Proficiency in scripting languages (Python, Bash).
  • Strong grasp of observability tools (Prometheus, Grafana, ELK, Sentry).
  • Good knowledge of cloud platforms (Huawei Cloud, Azure preferred)

Mandatory skills:

  • Strong background in Linux/Unix Administration
  • Solid hands-on experience deploying and operating Kubernetes or Openshift clusters
  • Experience configuring and maintaining monitoring and observability solutions
  • Ability to troubleshoot and resolve complex production issues efficiently, including performing root
  • cause analysis and restoring services quickly during high-pressure incidents or critical outages
  • Experience in backing up and restoring various systems
  • Working together with project managers and solution architects while serving as subject matter
  • Experts
  • Implementing basic network security (e.g. configuring VPCs, firewalls/security groups, etc.)
  • Understand the dependencies of various GPU cards, and upgrade container images as needed in
  • order to ensure compatibility
  • Deploy and operate products provided by third party providers
  • Creating releases together with the development team and deploying release packages to all required environments

Bonus Skills:

  • Good understanding of typical system architecture and interaction between its components
  • Experience automating tasks using infrastructure-as-code tools, e.g. Ansible, Terraform
  • Thorough understanding of a companys systems, including auxiliary components like caching
  • systems (e.g., Redis, Memcached) and message queues (e.g., RabbitMQ, Kafka)
  • Good understanding of databases, e.g. Postgres, Elasticsearch, Clickhouse
  • Basic scripting
  • Working knowledge of OAuth 2.0, OpenID/OpenID-Connect, SAML 2.0, Kerberos, LDAP
What working at AIQ offers:

Culture:Encouraging initiative, work in an environment that is fast-paced and varied, surrounded by talented peers from around the world, who are similarly attracted to applying their skills to solve transformational challenges.

Career:Join a team in which your contribution is recognized and rewarded, while you are supported to operate at your peak performance.

Rewards:An attractive renumeration package that includes healthcare, education support for dependents, leave benefits, and more.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Engineer - Site Reliability
Senior Engineer - Site Reliability

AIQ • Abu Dhabi

On-site
AED 300,000 - 480,000
Healthcare
Education support for dependents
Leave benefits
+1
Senior SRE: Reliability Platform Lead
Senior SRE: Reliability Platform Lead

MATRIXJV Co. Ltd. • Abu Dhabi

On-site
AED 300,000 - 540,000
Healthcare
Education support
Leave benefits
Senior Site Reliability Engineer – Abu Dhabi
Senior Site Reliability Engineer – Abu Dhabi

AIQ • Abu Dhabi

On-site
AED 300,000 - 480,000
Healthcare
Education support for dependents
Leave benefits
+1
Senior Specialist - Projects
Senior Specialist - Projects

Showcify • Abu Dhabi

On-site
AED 120,000 - 180,000
Healthcare
Education support for dependents
Leave benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether SRL • United Arab Emirates

On-site
AED 350,000 - 600,000
Fully remote
Global engineering org
Ownership over reliability
+2
Senior Specialist - Projects
Senior Specialist - Projects

AIQ • Abu Dhabi

On-site
AED 380,000 - 520,000
Healthcare
Education support for dependents
Leave benefits
Associate ML Ops Engineer
Associate ML Ops Engineer

Opus • Abu Dhabi

On-site
AED 120,000 - 180,000
21 days paid annual leave
Company health insurance
Visa sponsorship for international
Associate ML Ops Engineer
Associate ML Ops Engineer

Tanqeeb • Abu Dhabi

On-site
AED 120,000 - 240,000
Health insurance
Visa sponsorship
21 days annual leave
Associate ML Ops Engineer
Associate ML Ops Engineer

AppliedAI • Abu Dhabi

On-site
AED 156,000 - 234,000
Health insurance
Visa sponsorship
On-site Abu Dhabi HQ
Associate MLOps Engineer
Associate MLOps Engineer

AppliedAI • Abu Dhabi

On-site
AED 180,000 - 300,000
Health insurance
Visa sponsorship
21 days paid annual leave
+2