AI Automation QA

Intro Recruitment Asia

Philippines

On-site

PHP 558,000 - 893,000

Full time

28 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intro Recruitment Asia is seeking a QA and test automation professional to define and execute quality assurance strategies for AI-enabled workflows, with emphasis on reliability and safety.

You will build automated evaluation harnesses, contribute to incident response, and collaborate with developers to improve prompts, tools, and workflow designs.

The role focuses on monitoring production, managing dashboards, and ensuring compliance with privacy and regulatory requirements.

Qualifications

  • Minimum 3 years’ experience in QA, test automation, or DevOps roles (or 2 years with direct experience testing AI or ML-enabled systems).
  • Strong Python skills for test automation, evaluation harnesses, and basic data analysis.
  • Must have experience in Selenium/Playwright.
  • Must have experience in testing AI Behaviors in a system and not just using AI as an assistant.
  • Must have experience LLM testing e.g. Prompt & Response Testing, Hallucination & Safety Testing, RAG Testing.
  • High attention to detail, with a focus on issues that materially impact reliability and user trust.
  • Comfort working with evolving tools, frameworks, and testing practices.
  • Collaborative mindset, using evidence-based insights to influence product and engineering decisions.

Responsibilities

  • Definition and execution of testing and quality assurance strategies for AI-enabled workflows
  • Continuous evaluation and monitoring of system behavior in production environments
  • Contribution to auditability, risk management, and continuous quality improvement
  • Define quality criteria and testing strategies for agent workflows, covering accuracy, latency, safety, compliance, and operational risk
  • Build automated evaluation harnesses to assess agent performance, including hallucination rates, tool misuse, policy violations, and task success
  • Implement continuous production monitoring to detect anomalies, quality degradation, and emerging safety concerns
  • Develop and maintain automated test suites using Playwright for UI testing and custom scripts for API and workflow validation
  • Apply LLM evaluation frameworks to assess output quality, regression, and system drift over time
  • Produce and maintain dashboards and reports that communicate quality metrics and trends to engineering and stakeholders
  • Develop and maintain runbooks for common failure modes and contribute to incident response activities
  • Collaborate closely with developers to improve prompts, tool definitions, and workflow designs based on test results
  • Ensure testing, logging, and monitoring practices align with data privacy, audit, and regulatory requirements

Skills

QA
Test automation
DevOps
Python
Selenium
Playwright
AI testing
LLM testing
Attention to detail
Collaboration

Tools

Selenium
Playwright
Pytest
SQL
Prometheus
Grafana

Job description

  • Definition and execution of testing and quality assurance strategies for AI‑enabled workflows
  • Continuous evaluation and monitoring of system behavior in production environments
  • Contribution to auditability, risk management, and continuous quality improvement
  • Define quality criteria and testing strategies for agent workflows, covering accuracy, latency, safety, compliance, and operational risk
  • Build automated evaluation harnesses to assess agent performance, including hallucination rates, tool misuse, policy violations, and task success
  • Implement continuous production monitoring to detect anomalies, quality degradation, and emerging safety concerns
  • Develop and maintain automated test suites using Playwright for UI testing and custom scripts for API and workflow validation
  • Apply LLM evaluation frameworks to assess output quality, regression, and system drift over time
  • Produce and maintain dashboards and reports that communicate quality metrics and trends to engineering and stakeholders
  • Develop and maintain runbooks for common failure modes and contribute to incident response activities
  • Collaborate closely with developers to improve prompts, tool definitions, and workflow designs based on test results
  • Ensure testing, logging, and monitoring practices align with data privacy, audit, and regulatory requirements
Qualifications
Knowledge, Skills & Experience
  • Minimum 3 years’ experience in QA, test automation, or DevOps roles (or 2 years with direct experience testing AI or ML‑enabled systems)
  • Strong Python skills for test automation, evaluation harnesses, and basic data analysis
  • Must have experience in Selenium/Playwright
  • Must have experience in testing AI Behaviors in a system and not just using AI as an assistant.
  • Must have experience LLM testing e.g. Prompt & Response Testing, Hallucination & Safety Testing, RAG Testing.
  • High attention to detail, with a focus on issues that materially impact reliability and user trust
  • Comfort working with evolving tools, frameworks, and testing practices
  • Collaborative mindset, using evidence‑based insights to influence product and engineering decisions
Technical Skills (Required)
  • AI Evaluation: Deepeval, RAGAS, Evidently.AI (LLM quality, drift, and regression analysis)
  • Workflow Testing: API and agent workflow validation using custom scripts
  • Monitoring: Production quality monitoring and anomaly detection
  • Pytest or equivalent testing frameworks
  • SQL for querying logs, metrics, or evaluation datasets
  • Prometheus, Grafana, or similar monitoring tools
  • Familiarity with hallucination detection and AI safety patterns
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Application Tester (Night Shift) | 1x RTO/week
Agentic AI Application Tester (Night Shift) | 1x RTO/week

Comrise • Taguig

On-site
PHP 700,000 - 1,000,000
Agentic AI Application Tester
Agentic AI Application Tester

Willis Towers Watson • Manila

On-site
PHP 600,000 - 900,000
Agentic AI Application Tester
Agentic AI Application Tester

Willis Towers Watson • Taguig

On-site
AGENTIC AI APPLICATION TESTER
AGENTIC AI APPLICATION TESTER

Our Clients • Taguig

On-site
PHP 1,200,000 - 1,500,000
AGENTIC AI APPLICATION TESTER
AGENTIC AI APPLICATION TESTER

Create Synergies Inc. • Taguig

On-site
PHP 650,000 - 950,000
Senior AI Developer (Automators)
Senior AI Developer (Automators)

Luxoft • Mexico

On-site
PHP 1,200,000 - 1,800,000
AI Automation Tester (Hybrid Setup)
AI Automation Tester (Hybrid Setup)

blaseek • Manila

On-site
PHP 3,616,000 - 4,823,000
Agentic AI Developer
Agentic AI Developer

Intro Recruitment Asia • Philippines

On-site
PHP 1,200,000 - 2,400,000
Senior AI Automation Engineer
Senior AI Automation Engineer

Luxoft • Mexico

On-site
PHP 6,756,000 - 8,600,000
QA Lead (AI-Powered Testing)
QA Lead (AI-Powered Testing)

Indra Philippines, Inc. • Philippines

On-site
PHP 1,200,000 - 2,400,000