An application made for this job — a tailored resume and cover letter that speak straight to the posting.
AuraOne is seeking evaluators to remotely review bilingual generalists prompts and responses, scoring them against a versioned rubric. You will compare paired outputs, label edge cases, and draft structured feedback for retraining the modeling team.
Responsibilities include identifying hallucinations, instruction-following failures, and unsafe content with precise tags, while maintaining calibration against gold standards and ensuring consistent rubric application.
Bilingual Generalists is a remote evaluation track for reviewing bilingual generalists evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Category: Frontier Model Evaluation · Pay: $49–$98 / hr · Location: Remote — US-eligible · Contractor
Bilingual Generalists is a remote evaluation track for reviewing bilingual generalists evaluation prompts and responses against AuraOne's quality rubric.
Bilingual Generalists is a remote evaluation track for reviewing bilingual generalists evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
AI data reviewers help turn bilingual generalists evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Review frontier model outputs. Judge benchmark failures and calibrate other evaluators.
$49–$98 / hr
Expected arrangement: contractor , with program-defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation.
Creating a specialist profile records your experience and preferences. Starting role intake is a separate action that attaches this role to your candidate record.
Placement timing depends on program demand and reviewer confirmation.