Content Adversarial Red Team Analyst- English Philippines

Welo Data

Manila

On-site

PHP 864,000 - 1,296,000

Part time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Welo Data in the Philippines is seeking a Content Adversarial Red Team Analyst to test AI platforms against safety policies and compliance requirements. You will design and execute challenging test scenarios, developing diverse prompts and edge-case inputs to probe system behaviour.

The role is freelance with potential to convert to full-time, paying US$10 per hour, and requires strong written English, analytical thinking, and attention to detail.

Qualifications

  • Educational background or equivalent experience in Trust & Safety, Content Safety, Policy, Linguistics, Communications, Journalism, Research, AI Evaluation, or a related field.
  • Experience in content safety, trust and safety, AI evaluation, content moderation, policy enforcement, quality assurance, or related areas.
  • Strong understanding of content safety policies, policy enforcement, and common safety risks.
  • Strong understanding of written language, context, intent, and different ways users may communicate.
  • Ability to think creatively and develop challenging, unusual, or unexpected test scenarios.
  • Strong analytical and critical-thinking skills.
  • Ability to identify patterns, weaknesses, inconsistencies, and gaps in system behaviour.
  • Comfortable working with complex or ambiguous situations and making informed decisions based on defined requirements.
  • Ability to understand and consistently apply detailed testing guidelines and policies.
  • Strong written English comprehension and communication skills.
  • Strong documentation skills and attention to detail.
  • Familiarity with AI systems, large language models, adversarial testing, red teaming, or AI safety is preferred.

Responsibilities

  • Design and execute authorized adversarial test scenarios to evaluate AI systems against content safety policies and compliance requirements.
  • Develop diverse and challenging prompts, inputs, and scenarios to test system behaviour.
  • Explore edge cases, unusual inputs, and complex situations that may reveal gaps in safety controls or policy enforcement.
  • Evaluate AI responses and identify potential weaknesses, inconsistencies, or policy-related concerns.
  • Test how systems respond to different forms of language, context, intent, and user behaviour.
  • Identify recurring patterns or scenarios that may require further testing or improvement.
  • Document test scenarios, system responses, findings, and supporting evidence clearly.
  • Apply defined testing guidelines and project requirements consistently.
  • Review complex cases and use sound judgment when assessing system behaviour.
  • Maintain high levels of quality, accuracy, and attention to detail while completing assigned tasks.

Skills

Analytical thinking
Creative thinking
Attention to detail
Written English
Documentation skills
Problem solving

Education

Relevant field degree or equivalent experience

Tools

Adversarial testing methods

Job description

Job Description:

Job Overview

We are seeking creative and analytical Content Adversarial Red Team Analysts to test AI platforms and models against content safety policies and compliance requirements.

In this role, you will design and execute challenging test scenarios to identify gaps in how AI systems respond to complex, unusual, or unexpected inputs. You will explore edge cases and different approaches to assess whether platforms and models consistently follow defined safety policies and compliance boundaries.

An ideal candidate is curious, analytical, and able to think from different user perspectives. You should have a strong understanding of content safety and policy and be comfortable exploring complex scenarios to identify weaknesses, inconsistencies, or gaps in system behaviour.

Your work will help identify areas of improvement and support the development of safer and more reliable AI systems.

Project Details
  • Contract Type: Freelance, with the potential to convert to a full-time role.
  • Pay Rate: US$10 per hour
  • Location: Philippines
  • Language: English
Responsibilities
  • Design and execute authorized adversarial test scenarios to evaluate AI systems against content safety policies and compliance requirements.
  • Develop diverse and challenging prompts, inputs, and scenarios to test system behaviour.
  • Explore edge cases, unusual inputs, and complex situations that may reveal gaps in safety controls or policy enforcement.
  • Evaluate AI responses and identify potential weaknesses, inconsistencies, or policy-related concerns.
  • Test how systems respond to different forms of language, context, intent, and user behaviour.
  • Identify recurring patterns or scenarios that may require further testing or improvement.
  • Document test scenarios, system responses, findings, and supporting evidence clearly.
  • Apply defined testing guidelines and project requirements consistently.
  • Review complex cases and use sound judgment when assessing system behaviour.
  • Maintain high levels of quality, accuracy, and attention to detail while completing assigned tasks.
Required Qualifications
  • Educational background or equivalent experience in Trust & Safety, Content Safety, Policy, Linguistics, Communications, Journalism, Research, AI Evaluation, or a related field.
  • Experience in content safety, trust and safety, AI evaluation, content moderation, policy enforcement, quality assurance, or related areas.
  • Strong understanding of content safety policies, policy enforcement, and common safety risks.
  • Strong understanding of written language, context, intent, and different ways users may communicate.
  • Ability to think creatively and develop challenging, unusual, or unexpected test scenarios.
  • Strong analytical and critical-thinking skills.
  • Ability to identify patterns, weaknesses, inconsistencies, and gaps in system behaviour.
  • Comfortable working with complex or ambiguous situations and making informed decisions based on defined requirements.
  • Ability to understand and consistently apply detailed testing guidelines and policies.
  • Strong written English comprehension and communication skills.
  • Strong documentation skills and attention to detail.
  • Familiarity with AI systems, large language models, adversarial testing, red teaming, or AI safety is preferred.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Red Team Analyst: Content Adversarial Testing
AI Safety Red Team Analyst: Content Adversarial Testing

Welo Data • Manila

On-site
PHP 864,000 - 1,296,000
Freelancer - GenAI Experts
Freelancer - GenAI Experts

Alice • España

Hybrid
EUR 40,000 - 60,000
AI Red Team Specialist — Adversarial Testing (Remote)
AI Red Team Specialist — Adversarial Testing (Remote)

Mercor • Quezon City

Remote
PHP 3,652,000 - 7,304,000
Cybersecurity Threat Analyst - English Philippines
Cybersecurity Threat Analyst - English Philippines

Welo Data • Manila

On-site
PHP 691,000 - 864,000
DE033626-Delivery Operations Sr Analyst
DE033626-Delivery Operations Sr Analyst

Accenture in the Philippines • Quezon City

On-site
PHP 900,000 - 1,300,000
Elite GenAI Specialist: Adversarial & Prompt Design Remote
Elite GenAI Specialist: Adversarial & Prompt Design Remote

Alice • España

Hybrid
EUR 40,000 - 60,000
AI Adversarial Specialist - Fully Remote
AI Adversarial Specialist - Fully Remote

Mercor • Philippines

Remote
USD 90,000 - 150,000
Remote AI Adversarial Safety Specialist (English & Vietnamese)
Remote AI Adversarial Safety Specialist (English & Vietnamese)

Mercor • Philippines

On-site
PHP 1,439,000 - 2,117,000
LILT | Tagalog AI Content Experts
LILT | Tagalog AI Content Experts

LILT (Production) • Manila

Remote
PHP 42,388,000 - 101,730,000
Free access to NMT technology
Rapid payments via Tipalti
Opportunity to work on cutting-edge AI projects
+1
Trust & Safety Admin Analyst: Content Moderation & AI Governance
Trust & Safety Admin Analyst: Content Moderation & AI Governance

Accenture in the Philippines • Philippines

On-site
PHP 300,000 - 420,000