Principal AI Ops Engineer

Discovered MENA

Abu Dhabi

On-site

AED 350,000 - 650,000

Full time

28 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Discovered MENA is partnering with a leading Abu Dhabi organization to recruit a Principal AI Ops Engineer. The role focuses on production AI systems, spanning inference, model serving, deployment, observability, reliability and performance.

You will shape operational standards for safe, reliable and efficient AI deployments. The position targets hands-on leaders who can set standards while remaining close to engineering and production systems, with responsibility for cost, capacity and incident

Qualifications

  • Proven Staff/Principal level with production ML/LLM experience.
  • Hands-on GPU-based inference and model serving expertise.
  • Strong Python engineering for automation and tooling.

Responsibilities

  • Design, operate and optimise GPU-based inference and model-serving infra for production AI and LLM workloads.
  • Optimize model serving latency, throughput, batching and autoscaling while minimising cost.
  • Build automated release pipelines for models, prompts and configurations with canaries and rollbacks.
  • Establish AI observability across model, retrieval and orchestration layers.
  • Define and maintain SLOs, alerts and incident response for AI systems.
  • Own capacity planning and cost optimisation across GPU infrastructure.
  • Develop reusable deployment patterns and tooling for reliable AI shipping.
  • Lead complex production incidents, load testing and root-cause analysis.
  • Provide technical leadership and establish engineering standards for AI systems at scale.

Skills

GPU inference
Model serving
Kubernetes
Python automation
Observability
Cost optimisation

Tools

Terraform
Langfuse
LangSmith
Grafana
Prometheus
Arize Phoenix

Job description

We’re partnering with a leading organisation in Abu Dhabi that is building and operating advanced AI systems at significant scale.

They’re looking for a Principal AI Ops Engineer to take ownership of how AI and LLM systems operate in production, covering inference and model serving, deployment, observability, reliability and performance.

This is a Principal-level individual contributor role for someone who combines deep AI infrastructure expertise with strong software and reliability engineering fundamentals. You’ll set the operational standards that allow engineering teams to deploy and run production AI systems safely, reliably and efficiently.

Discover the Responsibilities:
  • Design, operate and optimise GPU-based inference and model-serving infrastructure for production AI and LLM workloads.
  • Optimise model serving across latency, throughput, batching, quantisation, autoscaling and infrastructure cost.
  • Build automated release pipelines for models, prompts and agent configurations, including canary deployments, regression gates and rollback strategies.
  • Establish AI-specific observability across model, retrieval and orchestration layers, including tracing, latency, cost and quality monitoring.
  • Define and maintain SLOs across availability, latency and AI system quality, alongside automated alerting and incident response processes.
  • Own capacity planning and cost optimisation across GPU and AI infrastructure.
  • Build secure, scalable Kubernetes environments and Infrastructure-as-Code patterns for production AI workloads.
  • Develop reusable deployment patterns, tooling and operational standards that enable engineering teams to ship AI systems reliably.
  • Lead complex production incidents, load testing and root-cause analysis across AI infrastructure and applications.
  • Provide technical leadership and help establish engineering standards for operating AI systems at scale.
Discover the Requirements:
  • Proven experience operating at Staff, Principal or equivalent senior IC level, with a track record of running production ML or LLM systems at scale.
  • Deep hands-on experience with GPU-based inference and model serving, including technologies such as vLLM, TGI, TensorRT-LLM or similar.
  • Strong understanding of batching, quantisation, latency, throughput, autoscaling and the performance trade-offs involved in production LLM serving.
  • Strong experience with AI/LLM observability, including tracing, quality monitoring, drift and regression detection.
  • Strong reliability engineering fundamentals across SLOs, incident response, capacity planning and post-mortems.
  • Strong Python engineering skills with experience building production-grade automation and infrastructure tooling.
  • Deep experience with Kubernetes, Docker and Infrastructure as Code, ideally Terraform, across cloud environments.
  • Experience with observability technologies such as Langfuse, LangSmith, Arize Phoenix, Grafana or Prometheus.
  • Experience integrating AI evaluation and regression testing into CI/CD and production release processes.
  • Experience with cloud infrastructure, ideally Azure, and production environments with strong security, data residency or compliance requirements.
  • Experience with GPU/AI infrastructure cost optimisation and FinOps would be advantageous.
  • A highly hands-on approach, with the ability to set technical standards while remaining close to engineering and production systems.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Ops Engineer
Principal AI Ops Engineer

Tanqeeb • Abu Dhabi

On-site
AED 320,000 - 480,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Client of Discovered MENA • Dubai

On-site
AED 350,000 - 520,000
Principal AI Ops Engineer: Production ML & LLM Reliability
Principal AI Ops Engineer: Production ML & LLM Reliability

Tanqeeb • Abu Dhabi

On-site
AED 320,000 - 480,000
AI/ML DevOps Specialist
AI/ML DevOps Specialist

Netision Technology LLP • United Arab Emirates

On-site
AED 320,000 - 520,000
Competitive salary
Cutting-edge AI/ML tech
Career growth opportunities
AI Delivery Manager
AI Delivery Manager

Quantum Talent Group • Dubai

On-site
AED 300,000 - 520,000
Senior Production GenAI Engineer
Senior Production GenAI Engineer

Tanqeeb • Dubai

On-site
AED 150,000 - 210,000
AI Ops Engineer (AI FinOps, Governance, Reliability, and Production Support)
AI Ops Engineer (AI FinOps, Governance, Reliability, and Production Support)

BlackCube Labs • Abu Dhabi

On-site
AED 240,000 - 360,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

GCS • Abu Dhabi

Hybrid
AED 250,000 - 420,000
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

EstateSight AI • Abu Dhabi

On-site
AED 450,000 - 900,000
Opportunity to lead AI innovation with societal impact
Research-driven environment
Leadership exposure across teams
AI Engineer (m/f/d)
AI Engineer (m/f/d)

Halian | Managed Services, Recruitment Agency & Contract Staffing • Abu Dhabi Emirate

On-site
AED 200,000 - 420,000