Stand out for this role — generate a tailored resume and cover letter in about a minute.
Tanqeeb in the United Arab Emirates seeks an AI Quality & Reliability Engineer to own Production Readiness of cloud-based AI solutions in a hybrid role, blending automated testing, AI model evaluation, and SRE.
You will design AI quality automation, implement chaos testing, CI/CD-integrated test suites in Python, define SLIs/SLOs, and perform security checks for strict banking data residency and privacy requirements.
We are seeking an AI Quality & Reliability Engineer to own the \"Production Readiness\" of our cloud-based AI solutions. This hybrid role combines automated software testing, AI model evaluation, and Site Reliability Engineering (SRE).
AI Quality Automation: Design and execute automated testing frameworks for AI services (e.g., Azure OpenAI, AWS Bedrock). This includes testing for model hallucinations, accuracy, and \"AI Content Safety\" latency.
Resiliency Engineering (SRE): Implement \"Chaos Engineering\" and load testing to ensure web/mobile backends can handle banking-scale traffic. Maintain high availability through automated recovery scripts.
Automated Regression: Build CI/CD-integrated test suites using Python that validate both the application logic and the infrastructure state (IaC validation).
Observability & SLIs: Define and monitor Service Level Indicators (SLIs) and Objectives (SLOs). Set up advanced alerting in Azure Monitor or AWS CloudWatch to catch performance degradation before users do.
Security & Compliance Testing: Automate security scans and compliance checks to ensure all AI data handling meets strict banking data residency and privacy protocols.
Technical & Professional Requirements:
Automation Stack: High proficiency in Python (for AI testing) and framework automation (PyTest, Selenium, or Robot Framework).
Cloud Infrastructure: Strong hands-on experience with Azure or AWS, specifically regarding networking, scaling, and serverless reliability.
AI/ML Understanding: Understanding of Prompt Engineering and how to evaluate AI model outputs (RAG evaluation, ROUGE/BLEU scores, or custom LLM-benchmarks).
Monitoring Tools: Experience with Grafana, Prometheus, or native cloud monitoring tools to build real-time reliability dashboards.
FinOps Awareness: Ability to identify \"expensive\" failing tests or inefficient cloud resource usage during the testing phase.
Languages: Python (Mandatory), Bash scripting.
Tools: GitHub Actions (CI/CD), Terraform (reading/validating), K6 or JMeter (Performance).
AI Frameworks: DeepEval, Ragas, or LangSmith (for automated AI evaluation).
Python
Azure/AWS
AI model evaluation