A leading evaluation firm in Canada is seeking an experienced evaluator to assess the accuracy and relevance of LLM-generated responses. The ideal candidate will have native-level fluency in French, strong English proficiency, and proven experience with large language models. Responsibilities include evaluating response quality, fact-checking, and creating structured feedback. Candidates should possess excellent writing skills and analytical thinking, and be able to work across various domains. This role offers competitive compensation and the opportunity to contribute to AI evaluations.
Qualifications
Native-level or near-native fluency in French with strong English proficiency.
Proven experience using large language models.
Excellent writing skills and ability to provide structured feedback.
Strong attention to detail and analytical thinking.
Ability to work across multiple topics and domains.
Background in structured analytical fields such as research, policy, analytics, linguistics, or engineering.
Responsibilities
Evaluate LLM-generated responses for accuracy, relevance, and effectiveness.
Perform fact-checking using reliable sources and tools.
Create high-quality evaluation data by annotating responses.
Assess reasoning quality, tone, clarity, and completeness of outputs.
Ensure responses follow expected conversational behavior and guidelines.
Apply consistent annotations using defined taxonomies, benchmarks, and evaluation frameworks.
Skills
Fluency in French
Writing skills
Attention to detail
Analytical thinking
Experience with large language models
Tools
External tools
Job description
Evaluate LLM-generated responses for accuracy, relevance, and effectiveness across a wide range of topics.
Perform fact-checking using reliable public sources and external tools.
Create high-quality human evaluation data by annotating response strengths, gaps, and factual errors.
Assess reasoning quality, tone, clarity, and completeness of AI-generated outputs.
Ensure responses follow expected conversational behavior and system guidelines.
Apply consistent annotations using defined taxonomies, benchmarks, and evaluation frameworks.
Requirements
Native-level or near-native fluency in French (ILR 5 / CEFR C2) with strong English proficiency.
Proven experience using large language models and understanding real-world LLM use cases.
Excellent writing skills with the ability to provide clear, nuanced, and structured feedback.
Strong attention to detail and analytical thinking.
Ability to work across multiple topics, domains, and evolving requirements.
Background in structured analytical fields such as research, policy, analytics, linguistics, or engineering.