AI Safety Red Team Evaluator

OpenTrain AI, Inc.

United Kingdom

Remote

GBP 12,000 - 24,000

Part time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote work
Self-set schedule
Flexible hours

Job summary

OpenTrain AI, Inc. is seeking a curious, entry-level Safety Validator to test how large language models respond to single-turn image-edit requests.

You will classify prompts and outputs according to project guidelines, review ambiguous or benign requests, and write precise rationales to support each decision. This is a part-time, contract engagement starting immediately, with 20+ hours weekly for 1–2 weeks.

Qualifications

  • Strong analytical judgment.
  • Ability to articulate clear rationales for classifications.
  • Experience or familiarity with AI safety concepts.

Responsibilities

  • Classify image-edit prompts and outputs per guidelines.
  • Review ambiguous, borderline, and benign requests consistently.
  • Write precise rationales supporting each classification decision.
  • Document bypassed safeguards and subtle policy violations.
  • Create adversarial prompts and compare model outputs.
  • Identify unclear or conflicting guidelines and suggest clarifications.

Skills

Policy analysis
Prompt classification
Rationale writing
AI safety concepts
Adversarial prompts

Education

Policy, law, ethics, linguistics, journalism, or computer science

Job description

The Work

You will test how large language models handle single-turn image-edit requests. You will classify prompts and outputs, find safety failures, and write clear explanations that can be used to improve model behavior.

  • Classify image-edit prompts and outputs using project guidelines and a defined safety taxonomy.
  • Review ambiguous, borderline, and benign requests consistently.
  • Write precise rationales that support each classification decision.
  • Document bypassed safeguards and subtle policy violations.
  • Create adversarial prompts and compare multiple model outputs.
  • Identify unclear or conflicting guidelines and suggest clarifications.
What It Pays And Takes

This is a project-based independent contractor engagement. The listing does not provide a pay rate. The role is marked entry level, but it requires strong analytical judgment and familiarity with AI safety concepts.

  • Hours: 20 or more hours per week.
  • Term: 1–2 weeks per statement of work.
  • Work arrangement: Part-time, independent contractor, with a self-set schedule.
  • Location: Open worldwide.
  • Language: Fluent English required.
  • Equipment: Your own desktop or laptop and a reliable internet connection.
  • Core skills: Policy-based analysis, careful classification, precise writing, and the ability to assess complex or ambiguous information.
  • Safety experience: Red teaming, prompt engineering, or designing challenge prompts to test AI safety filters.
  • Helpful background: Content moderation, policy analysis, AI safety evaluation, RLHF, or data annotation.
  • Relevant education or experience may include policy, law, ethics, linguistics, journalism, or computer science.
About AI Training Work

AI training work uses human judgments, examples, and written feedback to improve how artificial intelligence systems behave. Safety evaluators are paid for careful policy analysis because their decisions help identify harmful outputs and improve model responses.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Red Team Analyst for Model Evaluation
AI Safety Red Team Analyst for Model Evaluation

OpenTrain AI, Inc. • United Kingdom

Remote
GBP 12,000 - 24,000
Remote work
Self-set schedule
Flexible hours
AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Obsidian • Greater London

On-site
GBP 120,000 - 180,000
AI Safety Expert - Red Team
AI Safety Expert - Red Team

Mercor • Greater London

On-site
GBP 50,000 - 65,000
AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Mercor • Greater London

On-site
GBP 90,000 - 130,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
AI Safety Practitioner - Fully Remote
AI Safety Practitioner - Fully Remote

Mercor • Greater London

Remote
GBP 62,000 - 72,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • Greater London

On-site
GBP 70,000 - 110,000
English-Speaking AI Social-Context Response Evaluator
English-Speaking AI Social-Context Response Evaluator

OpenTrain AI, Inc. • United Kingdom

Remote
GBP 11,000 - 23,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • Greater London

On-site
GBP 70,000 - 110,000