Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
AuraOne is seeking a remote Trust and Safety Policy AI Evaluator to stress-test AI systems against adversarial prompts. You will craft attack scenarios, document failures, and map each jailbreak to the violated rubric clause to help patch gaps.
You will think like attackers, write rigorous failure reports, and help harden models before release. The role emphasizes compensation as an hourly contractor and eligibility from the US.
Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Category: AI Safety & Red Teaming · Pay: $65–$70 / hr · Location: Remote — US-eligible · Contractor
Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts.
Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them.
Push on refusal boundaries and dual-use risk before a model ships.
Track Adversarial evaluation Work model Remote · Independent specialist contractor Compensation Hourly rate confirmed after the interview process. Eligible from US
Hourly rate confirmed after the interview process.
Expected arrangement: contractor , with program-defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation.