Principal AI Ops Engineer

Tanqeeb

Abu Dhabi

On-site

AED 320,000 - 480,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Tanqeeb in Abu Dhabi is seeking a Principal AI Ops Engineer to own production AI systems, covering inference, deployment, observability, and reliability. You will set standards for efficient, safe AI operations at scale.

The role requires hands-on GPU-based inference expertise, Kubernetes, IaC, and strong Python engineering skills, with a focus on cost and performance optimization for enterprise AI workloads.

Qualifications

  • Proven senior IC with staff/principal level in production ML/LLM.
  • Hands-on GPU-based inference and model serving experience.
  • Deep knowledge of batching, quantisation, latency and throughput.
  • Strong AI/LLM observability, tracing and drift detection.
  • Reliability engineering: SLOs, incidents, capacity planning.
  • Hands-on Python for automation and tooling in production.
  • Kubernetes, Docker and IaC across cloud environments.
  • Experience with Azure and secure production environments.

Responsibilities

  • Design, operate and optimise GPU inference and model-serving infrastructure.
  • Optimise model serving latency, throughput, batching and costs.
  • Build automated release pipelines for models and configurations.
  • Establish AI observability across model, retrieval and orchestration layers.
  • Define and maintain SLOs with automated alerts and post-mortems.
  • Own capacity planning and cost optimisation for GPU/AI infra.
  • Build secure, scalable Kubernetes environments and IaC patterns.
  • Develop reusable deployment patterns and standards for AI systems.
  • Lead complex production incidents and root-cause analysis.
  • Provide technical leadership and engineering standards for AI systems.

Skills

GPU-based inference
Model serving
Python engineering
Kubernetes
Infrastructure as Code
Observability
Reliability engineering
FinOps awareness

Tools

Kubernetes
Docker
Terraform

Job description

Principal AI Ops Engineer – Abu Dhabi
Discover the Opportunity:

We’re partnering with a leading organisation in Abu Dhabi that is building and operating advanced AI systems at significant scale.

They’re looking for a Principal AI Ops Engineer to take ownership of how AI and LLM systems operate in production, covering inference and model serving, deployment, observability, reliability and performance.

This is a Principal-level individual contributor role for someone who combines deep AI infrastructure expertise with strong software and reliability engineering fundamentals. You’ll set the operational standards that allow engineering teams to deploy and run production AI systems safely, reliably and efficiently.

Discover the Responsibilities:
  • Design, operate and optimise GPU-based inference and model-serving infrastructure for production AI and LLM workloads.
  • Optimise model serving across latency, throughput, batching, quantisation, autoscaling and infrastructure cost.
  • Build automated release pipelines for models, prompts and agent configurations, including canary deployments, regression gates and rollback strategies.
  • Establish AI-specific observability across model, retrieval and orchestration layers, including tracing, latency, cost and quality monitoring.
  • Define and maintain SLOs across availability, latency and AI system quality, alongside automated alerting and incident response processes.
  • Own capacity planning and cost optimisation across GPU and AI infrastructure.
  • Build secure, scalable Kubernetes environments and Infrastructure-as-Code patterns for production AI workloads.
  • Develop reusable deployment patterns, tooling and operational standards that enable engineering teams to ship AI systems reliably.
  • Lead complex production incidents, load testing and root-cause analysis across AI infrastructure and applications.
  • Provide technical leadership and help establish engineering standards for operating AI systems at scale.
Discover the Requirements:
  • Proven experience operating at Staff, Principal or equivalent senior IC level, with a track record of running production ML or LLM systems at scale.
  • Deep hands-on experience with GPU-based inference and model serving, including technologies such as vLLM, TGI, TensorRT-LLM or similar.
  • Strong understanding of batching, quantisation, latency, throughput, autoscaling and the performance trade-offs involved in production LLM serving.
  • Strong experience with AI/LLM observability, including tracing, quality monitoring, drift and regression detection.
  • Strong reliability engineering fundamentals across SLOs, incident response, capacity planning and post-mortems.
  • Strong Python engineering skills with experience building production-grade automation and infrastructure tooling.
  • Deep experience with Kubernetes, Docker and Infrastructure as Code, ideally Terraform, across cloud environments.
  • Experience with observability technologies such as Langfuse, LangSmith, Arize Phoenix, Grafana or Prometheus.
  • Experience integrating AI evaluation and regression testing into CI/CD and production release processes.
  • Experience with cloud infrastructure, ideally Azure, and production environments with strong security, data residency or compliance requirements.
  • Experience with GPU/AI infrastructure cost optimisation and FinOps would be advantageous.
  • A highly hands-on approach, with the ability to set technical standards while remaining close to engineering and production systems.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Product Engineer
Principal AI Product Engineer

Discovered MENA • Abu Dhabi

On-site
AED 350,000 - 550,000
Principal AI Ops Engineer: Production ML & LLM Reliability
Principal AI Ops Engineer: Production ML & LLM Reliability

Tanqeeb • Abu Dhabi

On-site
AED 320,000 - 480,000
AI Ops Engineer (m/f/d)
AI Ops Engineer (m/f/d)

Halian | Managed Services, Recruitment and Contract Staffing Agency • Abu Dhabi

On-site
AED 220,000 - 380,000
AI Engineer (m/f/d)
AI Engineer (m/f/d)

Halian | Managed Services, Recruitment Agency & Contract Staffing • Abu Dhabi Emirate

On-site
AED 200,000 - 420,000
AI/ML DevOps Specialist
AI/ML DevOps Specialist

Netision Technology LLP • United Arab Emirates

On-site
AED 320,000 - 520,000
Competitive salary
Cutting-edge AI/ML tech
Career growth opportunities
Senior Engineer - Backend / AI
Senior Engineer - Backend / AI

Walker Lovell Ltd • Abu Dhabi

On-site
AED 350,000 - 500,000
Visa sponsorship for relocation
Opportunity for technical leadership
AI Ops Engineer
AI Ops Engineer

Tanqeeb • Abu Dhabi

On-site
AED 300,000 - 480,000
Principal AI Engineer
Principal AI Engineer

La Fosse • Abu Dhabi

Hybrid
AED 670,000 - 781,000
Tax free salary
Relocation package
Hybrid working
+1
AI Ops Engineer
AI Ops Engineer

D4 Insight • Abu Dhabi

On-site
AED 150,000 - 190,000
Agentic AI Engineer | Systems Ltd | Dubai, UAE
Agentic AI Engineer | Systems Ltd | Dubai, UAE

Systems Ltd • Dubai

On-site
AED 360,000 - 600,000