Remote | AI Evaluation Guidelines & Rubric Specialist — $40–$60/hour

24-Mag Llc

Northern (KY)

Hybrid

USD 55,000 - 83,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

24-MAG LLC is seeking experienced linguists, instructional designers, and technical writers for a remote, full-time role to develop rater guidelines and scoring rubrics for generative AI evaluation. You will translate complex requirements into actionable instructions for human evaluators across finance, retail, insurance, legal, sports, and more.

Ideal candidates have 3+ years in related fields and a track record of reducing ambiguity with clear guidance.

Qualifications

  • 3+ years of professional experience in linguistics, instructional design, technical writing, or related field.
  • Experience developing guidelines and rubrics for human evaluators in generative AI or RLHF projects.
  • Ability to resolve ambiguity and contradictions in complex specifications.
  • Translate specialist requirements into clear, practical instructions.
  • Ability to work across multiple domains with SME collaboration.

Responsibilities

  • Rater guideline development: translate requirements into clear, structured instructions.
  • Rubric & evaluation framework design with measurable criteria.
  • Ambiguity and consistency review to reduce contradictions.
  • Cross-domain instructional translation for finance, retail, insurance, legal, and more.

Skills

Linguistics
Instructional design
Technical writing
Content design
Guideline development
Rubric development
Evaluation design
RLHF awareness

Education

Linguistics or related degree
Instructional design or education degree
Graduate studies in applied linguistics/learning design
Equivalent professional experience in AI evaluation

Tools

Annotation platforms
Human-feedback workflows
Rater calibration processes
Documentation systems

Job description

We are sharing a specialised full-time consulting opportunity for US-based linguists, instructional designers, and technical writers experienced in developing clear evaluation guidelines, structured rubrics, and human-rating instructions for generative AI programmes.

This role supports a high-impact generative AI initiative focused on translating complex and potentially ambiguous programme requirements into precise, practical guidance for human evaluators. Selected professionals will develop rater-ready instructions across domains such as finance, retail, insurance, legal, and sports while resolving contradictions, defining edge cases, and improving consistency throughout evaluation workflows.

Key Responsibilities
Rater Guideline Development
  • Translate programme requirements into clear, structured, and actionable instructions for human evaluators
  • Develop guidelines that can be applied consistently across standard scenarios and complex edge cases
  • Define terminology, rating criteria, decision rules, exceptions, and escalation pathways
  • Ensure instructions are accessible to raters while preserving necessary domain-specific precision
Rubric & Evaluation Framework Design
  • Design detailed scoring rubrics for evaluating generative AI outputs
  • Establish measurable criteria covering correctness, relevance, reasoning quality, completeness, and instruction adherence
  • Create examples and counterexamples illustrating different performance levels
  • Align evaluation frameworks with programme objectives and quality standards
Ambiguity & Consistency Review
  • Review draft specifications for ambiguity, contradiction, missing information, and inconsistent terminology
  • Identify instructions that may lead to conflicting interpretations across raters
  • Revise guideline sets until they can be applied reliably with minimal escalation
  • Document concrete before-and-after improvements to written requirements and evaluation instructions
Cross-Domain Instructional Translation
  • Convert specifications from finance, retail, insurance, legal, sports, and other specialist domains into rater-ready guidance
  • Collaborate with subject matter experts to understand domain-specific terminology and professional judgment
  • Preserve important technical nuance while making instructions clear to non-specialist evaluators
  • Maintain consistent structure and quality across multiple domain-specific guideline sets
Ideal Profile
Strong candidates may have:
  • At least 3 years of professional experience in linguistics, instructional design, technical writing, content design, or a closely related field
  • Direct experience developing or refining guidelines and rubrics for human evaluators in generative AI, RLHF, or model-assessment programmes
  • Demonstrated ability to resolve ambiguity and contradiction in complex written specifications
  • Experience translating specialist requirements into clear and practical instructions
  • Ability to work effectively across multiple subject-matter domains
  • A portfolio or concrete examples showing measurable improvements to guidelines, rubrics, or instructional materials
  • Demonstrable professional growth and increasing responsibility
  • Reliable availability for at least 35 hours per week during weekdays
Educational Background
  • A degree in linguistics, instructional design, education, communications, technical writing, language studies, or a related field is highly relevant
  • Graduate-level education in applied linguistics, learning design, human-computer interaction, or information design may be helpful
  • Equivalent professional experience in AI evaluation, technical documentation, or guideline development may also be considered
  • Training in assessment design, taxonomy development, content strategy, or quality assurance may be valuable
Nice to Have
  • Experience supporting large language model evaluation, reinforcement learning from human feedback, or AI training-data programmes
  • Familiarity with annotation platforms, human-feedback workflows, and rater calibration processes
  • Experience developing domain-specific guidance for finance, insurance, retail, legal, sports, or comparable fields
  • Knowledge of controlled language, information architecture, taxonomy design, or content governance
  • Experience conducting guideline usability tests or analysing inter-rater consistency
  • Familiarity with version control, documentation systems, and structured authoring tools
  • Previous collaboration with researchers, programme managers, engineers, and subject matter experts
Why This Opportunity
  • Apply linguistic and instructional-design expertise to an advanced generative AI initiative
  • Influence the clarity and reliability of human evaluation processes
  • Develop guidelines used across a wide range of professional subject-matter domains
  • Solve complex problems involving ambiguity, edge cases, and evaluation consistency
  • Join a full-time remote engagement with competitive hourly compensation
Contract Details
  • Full-time W-2 contingent employment arrangement
  • Fully remote role available to candidates based in the United States
  • Expected commitment of at least 35 hours per week during weekdays
  • Competitive rates between $40–$60 per hour depending on expertise and project scope
  • Direct experience developing rater guidelines or rubrics for generative AI or RLHF programmes is required
  • Applicants should be prepared to provide concrete examples of guideline or specification improvements
  • Immediate availability is preferred
  • Work may include onboarding, calibration, documentation review, and ongoing guideline refinement
  • Project scope and duration may be adjusted according to programme requirements and performance
About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Rater Guidelines Writer (Linguist / Instructional Designer)
AI Rater Guidelines Writer (Linguist / Instructional Designer)

HumanitApp • Northern (KY)

Hybrid
USD 62,000 - 90,000
Remote | Senior Marketing Strategy & AI Evaluation Specialist — $55–$75/hour
Remote | Senior Marketing Strategy & AI Evaluation Specialist — $55–$75/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 76,000 - 103,000
Fully remote role
Competitive hourly compensation
Remote | Educator/ Assessor Writer — $90–$140/hour
Remote | Educator/ Assessor Writer — $90–$140/hour

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 124,000 - 193,000
Remote | Product Manager/Product Owner — $90–$140/hour
Remote | Product Manager/Product Owner — $90–$140/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 110,000 - 179,000
Remote | Senior Finance & AI Evaluation Specialist — $60–$85/hour
Remote | Senior Finance & AI Evaluation Specialist — $60–$85/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 83,000 - 117,000
Remote work
Competitive hourly rate
Remote AI Evaluation Guidelines Architect
Remote AI Evaluation Guidelines Architect

TrulyHired • United States

On-site
Confidential
Fully remote engagement
AI Evaluation Guidelines Architect
AI Evaluation Guidelines Architect

24-Mag Llc • Northern (KY)

Hybrid
USD 55,000 - 83,000
AI Rater Guidelines Writer (Linguist / Instructional Designer)
AI Rater Guidelines Writer (Linguist / Instructional Designer)

Mercor • United States

On-site
USD 90,000 - 120,000
AI Rater Guidelines Writer (Linguist / Instructional Designer)
AI Rater Guidelines Writer (Linguist / Instructional Designer)

DigiNo • Northern (KY)

Hybrid
USD 75,000 - 120,000
Generative AI Evaluator | $30/hr Remote
Generative AI Evaluator | $30/hr Remote

Crossing Hurdles • United States

Remote
Confidential