A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Job summary
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong analytical skills, attention to detail, and excellent communication abilities in English. This position allows for flexibility, with a commitment between 10 to 40 hours each week, compensated hourly at $20–$30.
Qualifications
Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical roles.
Proficiency in rubric-based scoring, benchmarking frameworks, and AI quality assessment.
Excellent attention to detail with strong decision-making skills in ambiguous cases.
Proficient English communication skills (written and verbal).
Ability to work independently in a remote environment.
Comfortable committing to structured evaluation workflows and evolving guidelines.
Responsibilities
Evaluate outputs from large language models and autonomous agent systems using defined rubrics and quality standards.
Review multi-step agent workflows, including screenshots and reasoning traces, to assess accuracy and completeness.
Apply benchmarking criteria consistently while identifying edge cases and recurring failure patterns.
Provide structured, actionable feedback to support model refinement and product improvements.
Participate in calibration sessions to ensure consistent evaluation alignment across reviewers.
Adapt to evolving guidelines and ambiguous scenarios with sound judgment.
Document findings clearly and communicate insights to relevant stakeholders.
Skills
LLM evaluation
AI output analysis
QA/testing
UX research
Attention to detail
Proficient English communication
Job description
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong analytical skills, attention to detail, and excellent communication abilities in English. This position allows for flexibility, with a commitment between 10 to 40 hours each week, compensated hourly at $20–$30.