Remote Bilingual AI Evaluation Specialist

AuraOne

United States

On-site

USD 28,000 - 48,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

AuraOne is seeking a Bilingual AI Response Evaluator to remotely review bilingual prompts and model outputs, applying AuraOne's quality rubric. You will compare paired responses, label edge cases, and write structured feedback for retraining efforts.

This contractor role requires strong attention to detail, clear reasoning, and the ability to work asynchronously for at least 10 hours per week. Prior experience in evaluation or annotation of bilingual content is preferred, with US remote

Qualifications

  • Prior evaluation, annotation, or human-rater experience on bilingual ai response evaluation or adjacent content for Bilingual AI Response Evaluator work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Responsibilities

  • Evaluate bilingual ai response evaluation model outputs against a versioned rubric and assign severity tags for Bilingual AI Response Evaluator assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Skills

Model output evaluation
Rubric-based annotation
Severity tagging
Inter-rater calibration
Bilingual AI Response evaluation
Voice, language and multimodal
AI evaluation
Rubric writing
Expert review

Job description

AuraOne is seeking a Bilingual AI Response Evaluator to remotely review bilingual prompts and model outputs, applying AuraOne's quality rubric. You will compare paired responses, label edge cases, and write structured feedback for retraining efforts.

This contractor role requires strong attention to detail, clear reasoning, and the ability to work asynchronously for at least 10 hours per week. Prior experience in evaluation or annotation of bilingual content is preferred, with US remote

Get your free, confidential resume review.
or drag and drop your file here.