AI Safety Evaluator for Real-World AI Systems

mpathic

Seattle (WA)

On-site

USD 70,000 - 110,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

mpathic.ai in Seattle is seeking AI Safety Experts for a temporary project to evaluate and improve the safety, reliability, and real-world behavior of frontier AI systems. This role suits professionals with expertise in human behavior, policy, education, healthcare, or trust & safety who can exercise sound judgment.

You will evaluate conversations, rate outputs with structured rubrics, annotate data for training and benchmarking, and identify risks and improvement opportunities while maintaining

Qualifications

  • Professional experience in psychology, behavioral science, trust & safety, or related research.
  • Strong written communication and attention to detail.
  • Able to learn structured evaluation frameworks and apply them consistently; strong critical thinking.
  • High ethical standards and sound judgment with sensitive content.
  • Willingness to sign NDAs and work on confidential projects.
  • Availability to work on-site in Seattle, San Francisco, or Boston.

Responsibilities

  • Evaluating AI-generated conversations, responses, and reasoning for quality, safety, and usefulness
  • Rating model outputs using structured evaluation rubrics and project guidelines
  • Annotating conversational data to support AI training and benchmarking
  • Identifying emerging risks, behavioral patterns, and opportunities for model improvement
  • Providing written feedback that helps researchers and engineers improve model performance
  • Maintaining strict confidentiality while working with proprietary AI systems and sensitive content
  • Participating in calibration sessions and quality reviews to ensure consistent evaluations

Skills

Psychology
Behavioral science
Trust & Safety
Research
Critical thinking
Written communication

Tools

Google Workspace
Slack

Job description

mpathic.ai in Seattle is seeking AI Safety Experts for a temporary project to evaluate and improve the safety, reliability, and real-world behavior of frontier AI systems. This role suits professionals with expertise in human behavior, policy, education, healthcare, or trust & safety who can exercise sound judgment.

You will evaluate conversations, rate outputs with structured rubrics, annotate data for training and benchmarking, and identify risks and improvement opportunities while maintaining

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Expert (Seattle or Boston)
AI Safety Expert (Seattle or Boston)

mpathic • Seattle (WA)

On-site
USD 70,000 - 110,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Evaluator & Model Alignment Expert
AI Safety Evaluator & Model Alignment Expert

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • New York (NY)

On-site
USD 120,000 - 190,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
null
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • New York (NY)

On-site
USD 120,000 - 170,000