AI Safety Red Team Analyst: Content Adversarial Testing

Welo Data

Manila

On-site

PHP 864,000 - 1,296,000

Part time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Welo Data in the Philippines is seeking a Content Adversarial Red Team Analyst to test AI platforms against safety policies and compliance requirements. You will design and execute challenging test scenarios, developing diverse prompts and edge-case inputs to probe system behaviour.

The role is freelance with potential to convert to full-time, paying US$10 per hour, and requires strong written English, analytical thinking, and attention to detail.

Qualifications

  • Educational background or equivalent experience in Trust & Safety, Content Safety, Policy, Linguistics, Communications, Journalism, Research, AI Evaluation, or a related field.
  • Experience in content safety, trust and safety, AI evaluation, content moderation, policy enforcement, quality assurance, or related areas.
  • Strong understanding of content safety policies, policy enforcement, and common safety risks.
  • Strong understanding of written language, context, intent, and different ways users may communicate.
  • Ability to think creatively and develop challenging, unusual, or unexpected test scenarios.
  • Strong analytical and critical-thinking skills.
  • Ability to identify patterns, weaknesses, inconsistencies, and gaps in system behaviour.
  • Comfortable working with complex or ambiguous situations and making informed decisions based on defined requirements.
  • Ability to understand and consistently apply detailed testing guidelines and policies.
  • Strong written English comprehension and communication skills.
  • Strong documentation skills and attention to detail.
  • Familiarity with AI systems, large language models, adversarial testing, red teaming, or AI safety is preferred.

Responsibilities

  • Design and execute authorized adversarial test scenarios to evaluate AI systems against content safety policies and compliance requirements.
  • Develop diverse and challenging prompts, inputs, and scenarios to test system behaviour.
  • Explore edge cases, unusual inputs, and complex situations that may reveal gaps in safety controls or policy enforcement.
  • Evaluate AI responses and identify potential weaknesses, inconsistencies, or policy-related concerns.
  • Test how systems respond to different forms of language, context, intent, and user behaviour.
  • Identify recurring patterns or scenarios that may require further testing or improvement.
  • Document test scenarios, system responses, findings, and supporting evidence clearly.
  • Apply defined testing guidelines and project requirements consistently.
  • Review complex cases and use sound judgment when assessing system behaviour.
  • Maintain high levels of quality, accuracy, and attention to detail while completing assigned tasks.

Skills

Analytical thinking
Creative thinking
Attention to detail
Written English
Documentation skills
Problem solving

Education

Relevant field degree or equivalent experience

Tools

Adversarial testing methods

Job description

Welo Data in the Philippines is seeking a Content Adversarial Red Team Analyst to test AI platforms against safety policies and compliance requirements. You will design and execute challenging test scenarios, developing diverse prompts and edge-case inputs to probe system behaviour.

The role is freelance with potential to convert to full-time, paying US$10 per hour, and requires strong written English, analytical thinking, and attention to detail.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Content Adversarial Red Team Analyst- English Philippines
Content Adversarial Red Team Analyst- English Philippines

Welo Data • Manila

On-site
PHP 864,000 - 1,296,000
AI Red Team Specialist — Adversarial Testing (Remote)
AI Red Team Specialist — Adversarial Testing (Remote)

Mercor • Quezon City

Remote
PHP 3,652,000 - 7,304,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • Quezon City

On-site
PHP 600,000 - 1,200,000
Remote AI Safety & Adversarial Red Team Expert
Remote AI Safety & Adversarial Red Team Expert

Mercor • Philippines

On-site
PHP 4,316,000 - 6,782,000
Remote AI Adversarial Safety Specialist (English & Vietnamese)
Remote AI Adversarial Safety Specialist (English & Vietnamese)

Mercor • Philippines

On-site
PHP 1,439,000 - 2,117,000
Remote AI Safety Red Team Engineer
Remote AI Safety Red Team Engineer

Mercor • Philippines

On-site
PHP 600,000 - 900,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Obsidian • Quezon City

On-site
PHP 5,478,000 - 9,130,000
AI Security Engineer: Advanced Threat Defense for AI Agents
AI Security Engineer: Advanced Threat Defense for AI Agents

Coins • Hinoba-an

Hybrid
PHP 900,000 - 1,400,000
AI Adversarial Specialist - Fully Remote
AI Adversarial Specialist - Fully Remote

Mercor • Philippines

Remote
USD 90,000 - 150,000
Remote AI Adversarial Tester (English/Vietnamese)
Remote AI Adversarial Tester (English/Vietnamese)

Mercor • Philippines

Remote
USD 90,000 - 150,000