QA Automation Engineer

GCS Recruitment Specialists

Abu Dhabi

On-site

AED 240,000 - 420,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GCS Recruitment Specialists in Abu Dhabi is seeking a Senior QA Automation Engineer to lead testing of agentic AI systems. You will design and run frameworks that validate complex models, drawings, and regulatory documents, ensuring trust and compliance in automated decisions.

You will own testing across integration, API, and performance, focusing on non-deterministic AI behaviors, grounding, and tool-use verification, with clear reporting and modern QE practices.

Qualifications

  • 5+ years of QA automation or quality engineering experience.
  • Experience testing ML models, LLM apps, or AI agents.
  • Independent operator who can scope and build testing frameworks.

Responsibilities

  • Agentic System Testing: evaluate tool selection, multi-step trajectories, and recovery from failures.
  • Grounding and citation verification to ensure outputs are supported by sources.
  • Honest-failure and refusal calibration to avoid unwanted outputs and refusals.
  • Guardrail and adversarial testing for prompt injections and untrusted data handling.
  • Multi-turn and state validation to ensure continuity across sessions.
  • AI & LLM validation: develop regression suites and judge outputs using gold standards.
  • Live production quality: monitor and evaluate live outputs and trigger alerts for regressions.

Skills

Python
LLM evaluation
Agent tracing
CI/CD
SQL data validation
Performance testing
Test automation libraries
Version control
QA reporting

Job description

QA Automation Engineer, AI & Agentic Systems

Level: Senior IC
Location: Abu Dhabi

Role Summary

We are building agentic AI systems that read complex models, drawings, specifications, and regulatory documents to check compliance, plan and validate schedules, and support design.

This role exists to make those agents trustworthy enough to act on.

This is a senior, hands-on quality role weighted toward testing non-deterministic AI and agentic systems. That is where most of your time sits and where the hiring bar is highest. You will also own the broader quality surface: integration, API, and performance testing are part of the remit, not out of scope.

The differentiator we are hiring for is the ability to define what "good" looks like when a system reasons, calls tools, and can be wrong in subtle ways.

You will scope, build, and run the frameworks yourself, with the independence of a senior engineer, and decide what to build first.

Key Responsibilities
  1. Agentic System Testing (primary focus)
  • Tool-Use & Trajectory Evaluation: Test whether agents select the right tools for the right reasons and follow sound multi-step trajectories, not just whether the final answer looks plausible. Evaluate planning, intermediate steps, and recovery when a tool fails or returns nothing.
  • Grounding & Citation Verification: Verify that agent claims are backed by cited evidence in the source model, drawing, or document, and that a stated location actually supports the answer, catching confident but unsupported outputs.
  • Honest-Failure & Refusal Calibration: Assert that agents ask for clarification or state when something cannot be determined instead of inventing an answer, and build tests where the correct behaviour is refusal.
  • Guardrail & Adversarial Testing: Probe prompt injection, jailbreaks, and instructions hidden inside ingested documents, ensuring the agent treats source content as untrusted data.
  • Multi-Turn & State: Validate follow-ups, references to prior turns, and that conversational state carries correctly across a session.
  1. AI & LLM Validation
  • Non-Deterministic Testing: Architect automated frameworks that score generative-AI outputs for hallucination, consistency, and factual accuracy against gold-standard datasets, using LLM-as-judge methods calibrated against human judgement.
  • Prompt & Model Regression: Design regression suites that catch prompt drift and model version drift, so changes to models or system instructions do not quietly degrade quality. Own the ground-truth and evaluation datasets these depend on.
  1. Live Production Quality
  • Continuous Evaluation: Extend evaluation beyond pre-release into production, continuously scoring live agent outputs so quality is measured on real traffic, not only in the test environment.
  • Monitoring & Alerting: Build quality monitoring that flags regressions, drift, and anomalous agent behaviour in production before users or customers do.
  • Quality Incident Response: Triage quality incidents and close the loop, tracing failures in AI logic back to the specific model version or dataset that caused them and feeding fixes into the development cycle.
  1. Integration, API & Performance
  • Backend, UI & API Testing: Build robust integration tests that validate API integration across services and key user-facing flows.
  • Secure Gateway Validation: Automate testing of secure API gateways, verifying that role-based access controls and sensitive-data redaction logic work correctly before data reaches AI models.
  • Performance & Load: Own performance test plans and their implementation using appropriate performance and load-testing technologies, validating latency, throughput, and stability under realistic load.
  1. Data, Traceability & Quality Gates
  • Data Validation: Use SQL and data-validation tooling to verify data quality across data platforms and vector databases, including the ground-truth and retrieval corpora the agents depend on.
  • Requirements Traceability: Map test and evaluation cases to system requirements and user needs, producing the verification-and-validation evidence and quality reports needed to ship with confidence.
  • Quality Gates: Enforce quality gates in CI/CD pipelines that prevent non-compliant models or code from merging and prepare readiness evidence for stage and release reviews.
Technical Requirements
  • AI Evaluation (core): Hands-on experience with LLM and agent-evaluation frameworks or custom Python evaluators, including LLM-as-judge techniques.
  • Agent Observability: Experience tracing and debugging agent runs, including tool calls, intermediate steps, token usage, and latency, using modern agent-observability and tracing technologies.
  • Core Automation: Expert Python skills for custom test harnesses and evaluation tooling, plus experience with standard UI and API automation libraries.
  • Performance Testing: Proven ability to craft and implement performance test plans using modern load and performance-testing tools.
  • Data Validation: Proficiency with SQL and data-validation tools, with familiarity with vector databases and retrieval corpora.
  • CI/CD Integration: Experience integrating automated tests and evaluations into CI/CD pipelines and enforcing quality gates.
  • Test Management & Reporting: Experience managing test reports and artifacts and communicating results clearly.
  • Version Control & QE Practices: Experience maintaining code-based frameworks in version control and applying modern quality-engineering practices for fast-paced teams, including shift-left testing, the test pyramid, and automation.
  • Traceability Tools: Familiarity with requirements-management tools and linking results to requirement IDs.
Professional Qualifications
  • Experience: 5+ years in QA automation or quality engineering, with at least 2 years focused on testing ML models, LLM applications, or AI agents.
  • Probabilistic-Systems Judgement: Able to define pass/fail criteria for systems whose outputs are not identical every run and to communicate confidence levels clearly to engineering leadership.
  • Independent Operator: A senior doer who scopes and builds testing and evaluation frameworks with minimal direction and prioritises what matters first.
  • Collaboration: Works closely with engineering and product, understands existing systems quickly, and moves fast.
Nice to Have
  • Industry Domain Experience: Familiarity with complex technical data and workflows, including models, drawings, specifications, compliance, and scheduling, is a strong plus.
  • Formal V&V Exposure: Exposure to structured systems-engineering governance, stage-gate reviews, and formal verification and validation is welcome but not required. We value the discipline more than the certification.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Junior AI Engineer
Junior AI Engineer

mindX360 • Dubai

On-site
AED 210,000 - 360,000
AI Engineer / Agent Developer
AI Engineer / Agent Developer

Al Gurg Group • Dubai

On-site
AED 350,000 - 650,000
Quality Assurance Automation Engineer
Quality Assurance Automation Engineer

Atain • Abu Dhabi Emirate

On-site
AED 360,000 - 600,000
QA Tester - AI
QA Tester - AI

Dicetek LLC • Abu Dhabi

On-site
AED 300,000 - 420,000
Senior Quality Assurance Engineer
Senior Quality Assurance Engineer

Institute of Foundation Models • Abu Dhabi

On-site
AED 167,400 - 223,200
AI Quality & Reliability Engineer (QA/SRE)
AI Quality & Reliability Engineer (QA/SRE)

Tanqeeb • Abu Dhabi

Hybrid
AED 180,000 - 260,000
QA Analyst 1 (UAE Nationals)
QA Analyst 1 (UAE Nationals)

noon • Dubai

On-site
AED 134,000 - 223,000
AI Quality & Reliability Engineer (QA/SRE)
AI Quality & Reliability Engineer (QA/SRE)

Dicetek LLC • Dubai

On-site
AED 201,000 - 290,000
AI Quality & Reliability Engineer – QA/SRE
AI Quality & Reliability Engineer – QA/SRE

D4Insight • Abu Dhabi

Hybrid
AED 280,000 - 420,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Client of Discovered MENA • Dubai

On-site
AED 350,000 - 520,000