AI Model Policy Trainer, Content Risk

Jobtailor

Seattle (WA)

On-site

USD 65,000 - 90,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Jobtailor seeks a policy evaluation professional to assess AI model responses involving violence and related content. You will determine classifications in context, distinguishing fiction, education, history, and defensive uses from real-world uplift or harm.

Responsibilities include crafting concise rationales, refining prompts, and participating in calibration discussions to align with policy language. Strong written communication and attention to detail are essential.

Qualifications

  • Strong judgment about violence in fiction, the real world, or among people in distress.
  • Ability to distinguish fictional, educational, historical, and defensive violence from real-world uplift or intent to harm.
  • Ability to assess meaningful real-world capability in AI model responses.
  • Ability to distinguish anger, frustration, or dark humor from credible threats or crisis indicators.
  • Ability to write concise rationales citing policy language and conversation details.
  • Ability to write and refine adversarial or borderline prompts.
  • Ability to identify policy gaps, contradictions, and emerging edge cases.
  • Ability to participate in calibration discussions and apply customer policy consistently.
  • Accuracy and attention to detail during repetitive, feedback-heavy evaluations.
  • Clear and precise written communication.
  • Ability to engage carefully, responsibly, and sustainably with graphic material.

Responsibilities

  • Evaluate user requests and AI model responses involving violence, weapons, threats, and dark fiction in full conversation context
  • Distinguish fictional, educational, historical, and defensive violence from requests seeking real-world uplift or expressing real intent to harm
  • Assess whether model responses provide meaningful real-world capability
  • Distinguish anger, frustration, and dark humor from credible threats or crisis indicators
  • Select defensible classifications for ambiguous cases and write concise policy-based rationales
  • Write and refine adversarial or borderline prompts
  • Identify policy gaps, contradictions, and emerging edge cases and raise them with project leads and policy teams
  • Participate in calibration discussions and update judgments based on stronger reasoning
  • Apply customer policy consistently without substituting personal beliefs
  • Maintain accuracy and attention to detail across repeated evaluations involving graphic material

Skills

Policy evaluation
Strong judgment
Written rationale
Attention to detail
Crisis indicators
Adversarial prompts
Calibration discussions
Policy gaps

Job description

  • Evaluate user requests and AI model responses involving violence, weapons, threats, and dark fiction in full conversation context
  • Distinguish fictional, educational, historical, and defensive violence from requests seeking real-world uplift or expressing real intent to harm
  • Assess whether model responses provide meaningful real-world capability
  • Distinguish anger, frustration, and dark humor from credible threats or crisis indicators
  • Select defensible classifications for ambiguous cases and write concise policy-based rationales
  • Write and refine adversarial or borderline prompts
  • Identify policy gaps, contradictions, and emerging edge cases and raise them with project leads and policy teams
  • Participate in calibration discussions and update judgments based on stronger reasoning
  • Apply customer policy consistently without substituting personal beliefs
  • Maintain accuracy and attention to detail across repeated evaluations involving graphic material
Requirements
  • Strong judgment about violence in fiction, the real world, or among people in distress
  • Ability to distinguish fictional, educational, historical, and defensive violence from real-world uplift or intent to harm
  • Ability to assess meaningful real-world capability in AI model responses
  • Ability to distinguish anger, frustration, or dark humor from credible threats or crisis indicators
  • Ability to write concise rationales citing policy language and conversation details
  • Ability to write and refine adversarial or borderline prompts
  • Ability to identify policy gaps, contradictions, and emerging edge cases
  • Ability to participate in calibration discussions and apply customer policy consistently
  • Accuracy and attention to detail during repetitive, feedback-heavy evaluations
  • Clear and precise written communication
  • Ability to engage carefully, responsibly, and sustainably with graphic material
  • A degree, a clearance, and a technical background are not required
  • Must be authorized to work lawfully in the United States for Handshake
Core Competencies

Demonstrates strong judgment and analytical skills in evaluating AI model responses related to violence and threats, while maintaining accuracy and attention to detail. Capable of writing concise rationales and engaging with graphic material responsibly.

Highest-signal resume keywords
  • Strong Judgment About Violence
  • Ability To Distinguish Fictional And Real-World Violence
  • Ability To Write Concise Rationales
  • Ability To Identify Policy Gaps
  • Clear And Precise Written Communication
Soft Skills
  • Attention To Detail
  • Analytical Skills
  • Engagement With Graphic Material
Industry Keywords
  • AI Model Evaluation
  • Policy-Based Rationales
  • Calibration Discussions
  • Crisis Indicators
  • Adversarial Prompts
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Policy Evaluator, Violence & Threats | Seattle Onsite
AI Safety Policy Evaluator, Violence & Threats | Seattle Onsite

Handshake • Seattle (WA)

On-site
USD 76,000 - 124,000
Benefits eligible
AI Safety Policy Evaluator, Violence & Threats | Seattle Onsite
AI Safety Policy Evaluator, Violence & Threats | Seattle Onsite

Apply • Seattle (WA), Northern (KY)

Hybrid
USD 76,000 - 124,000
AI Safety Policy Evaluator, Violence & Threats | Remote US
AI Safety Policy Evaluator, Violence & Threats | Remote US

Apply • Northern (KY)

Hybrid
USD 62,000 - 76,000
Benefits eligible
AI Model Policy Trainer, Generalist (Seattle)
AI Model Policy Trainer, Generalist (Seattle)

Cacheflow • Seattle (WA)

On-site
USD 48,000 - 138,000
AI Content Risk Policy Trainer: Violence & Safety
AI Content Risk Policy Trainer: Violence & Safety

Jobtailor • Seattle (WA)

On-site
USD 65,000 - 90,000
AI Violence & Safety Policy Evaluator
AI Violence & Safety Policy Evaluator

handshake • United States

Remote
USD 90,000 - 150,000
AI Model Policy Trainer, Mental Health
AI Model Policy Trainer, Mental Health

Cacheflow • Seattle (WA)

On-site
USD 110,000 - 140,000
AI Policy Generalist - Seattle Onsite
AI Policy Generalist - Seattle Onsite

Handshake • Seattle (WA)

On-site
USD 62,000 - 76,000
AI Policy Generalist - Seattle Onsite
AI Policy Generalist - Seattle Onsite

Apply • Seattle (WA)

On-site
USD 62,000 - 76,000
AI Violence & Fiction Policy Specialist
AI Violence & Fiction Policy Specialist

Apply • Seattle (WA), Northern (KY)

Hybrid
USD 76,000 - 124,000