AI Safety Specialist - Evaluation Expert

Visa Hunt

España

On-site

PHP 5,182,000 - 6,046,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor connects elite creative and technical talent with leading AI research labs. The AI Safety Practitioner role is a remote contract position focused on evaluating AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.

The candidate pool should include a Bachelor's degree or higher and at least 5 years in AI Safety or related fields, with excellent written English and strong analytical reasoning, capable of assessing nuanced, policy-sensitive scenarios

Qualifications

  • Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or related discipline.
  • 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or related field.
  • Excellent written English, critical thinking, and analytical reasoning skills.

Responsibilities

  • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
  • Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains.
  • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
  • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
  • Provide structured feedback to improve model alignment and safety performance.
  • Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.

Skills

Excellent written English
Critical thinking
Analytical reasoning
Policy-sensitive evaluation
5+ years AI Safety experience

Education

Bachelor's degree or higher in related discipline

Job description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Position: AI Safety Practitioner
Type:Contract
Compensation:$60–$70/hour
Location:Remote

Role Responsibilities
  • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
  • Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains.
  • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
  • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
  • Provide structured feedback to improve model alignment and safety performance.
  • Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
Qualifications

Must-Have

  • Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
  • 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field.
  • Excellent written English, critical thinking, and analytical reasoning skills.
  • Ability to consistently evaluate nuanced and policy-sensitive scenarios.
Preferred
  • Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation.
  • Familiarity with safety policies, content moderation, or evaluation rubric development.
  • Experience reviewing complex, high-risk, or ambiguous content.
Resources & Support
  • For details about the interview process and platform information, please check:
  • For any help or support, reach out to:

Originally posted on Himalayas

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Safety Evaluator & Benchmark Expert
Remote AI Safety Evaluator & Benchmark Expert

Visa Hunt • España

On-site
PHP 5,182,000 - 6,046,000
Remote AI Safety Specialist — RLHF & Alignment
Remote AI Safety Specialist — RLHF & Alignment

Visa Hunt • España

Remote
PHP 5,182,000 - 6,046,000
AI Safety & Policy Expert - English Speakers
AI Safety & Policy Expert - English Speakers

TELUS Digital AI Data Solutions • Philippines

On-site
PHP 1,554,000 - 4,316,000
Elite AI CBRN Specialist — Remote Freelance
Elite AI CBRN Specialist — Remote Freelance

Alice • España

Hybrid
EUR 50,000 - 80,000
Academic Evaluator - Fully Remote | Upto $160/hr
Academic Evaluator - Fully Remote | Upto $160/hr

Visa Hunt • España

Remote
PHP 6,910,000 - 13,819,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • Quezon City

On-site
PHP 600,000 - 1,200,000
Content Adversarial Red Team Analyst- English Philippines
Content Adversarial Red Team Analyst- English Philippines

Welo Data • Manila

On-site
PHP 864,000 - 1,296,000
Freelancer - AI CBRN Experts
Freelancer - AI CBRN Experts

Alice • España

Hybrid
EUR 50,000 - 80,000
AI Safety & Content Policy Specialist (Remote)
AI Safety & Content Policy Specialist (Remote)

TELUS Digital AI Data Solutions • Manila

On-site
PHP 1,726,000 - 3,576,000
Remote AI Safety Red Teamer - English & Vietnamese
Remote AI Safety Red Teamer - English & Vietnamese

Visa Hunt • Philippines

Remote
PHP 1,468,000 - 2,159,000