AI/ML Evaluator - English

Welocalize

Town of Texas (WI)

On-site

USD 192,864,000 - 220,416,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Welocalize is seeking analytical and technically skilled AI/ML Evaluators to review and evaluate complex AI system behaviour using expert human judgment. You will analyse AI system outputs, telemetry and other technical signals to identify meaningful patterns and distinguish legitimate user activity from automated or bot activity.

You'll provide high-quality human evaluations to train, evaluate and improve AI systems, applying your knowledge of AI/ML concepts and detailed project guidelines with

Qualifications

  • Educational background or equivalent experience in AI, ML, CS, data science, statistics, engineering, or related field.
  • Experience with AI/ML systems, data analysis, or technical system analysis.
  • Strong understanding of core AI/ML concepts.
  • Excellent written English and documentation skills.
  • Familiarity with bot detection, anomaly detection, or fraud detection is preferred.
  • Experience with data annotation or model evaluation is preferred.

Responsibilities

  • Review and evaluate AI system behaviour, outputs, and technical data according to project guidelines.
  • Analyse system telemetry, event data, signals, and other technical information to identify meaningful patterns.
  • Evaluate patterns and signals that distinguish legitimate user activity from automated or bot activity.
  • Apply AI/ML knowledge and expert human judgment when reviewing complex or ambiguous cases.
  • Identify unusual patterns, inconsistencies, or behaviours that may require further review.
  • Evaluate complex cases using available evidence, context, and defined project requirements.
  • Classify, label, or annotate assigned data accurately and consistently.
  • Provide high-quality human evaluations to support AI system training, evaluation, and improvement.
  • Document decisions and supporting reasoning clearly when required.
  • Identify unclear or unusual cases and flag them according to defined project processes.
  • Apply detailed project guidelines consistently across assigned tasks.
  • Maintain high levels of quality, accuracy, consistency, and attention to detail.

Skills

Analytical thinking
AI/ML concepts
Data analysis
Pattern recognition
Documentation skills

Education

Bachelor's degree or equivalent in AI/ML/CS/Data Science/Statistics/Engineering

Job description

Job Overview

We are seeking analytical and technically skilled AI/ML Evaluators to review and evaluate complex AI system behaviour using expert human judgment.

In this role, you will analyse AI system outputs, system telemetry, and other relevant technical signals to evaluate patterns of behaviour. A key part of the role will involve helping distinguish legitimate user activity from sophisticated automated or bot activity based on available data and defined project guidelines.

Your evaluations will provide high-quality, human-verified data used to train, evaluate, and improve AI systems. You will work with complex and sometimes ambiguous information, applying your understanding of AI/ML concepts and analytical reasoning to make accurate and consistent judgments.

An ideal candidate has a strong understanding of AI and machine learning concepts and is comfortable analysing technical data and system behaviour. You should be able to identify meaningful patterns, interpret complex signals, and make informed decisions based on evidence and project requirements.

Your work will help improve the accuracy and reliability of AI systems used to identify automated abuse while reducing the risk of incorrectly classifying legitimate user activity.

Project Details
  • Contract Type: Freelance, with the potential to convert to a full-time role.
  • Pay Rate: US$72 per hour
  • Location: United States
  • Language: English
Responsibilities
  • Review and evaluate AI system behaviour, outputs, and relevant technical data according to defined project guidelines.
  • Analyse system telemetry, event data, signals, and other technical information to identify meaningful patterns.
  • Evaluate patterns and signals that may help distinguish legitimate user activity from automated or bot activity.
  • Apply AI/ML knowledge and expert human judgment when reviewing complex or ambiguous cases.
  • Identify unusual patterns, inconsistencies, or behaviours that may require further review.
  • Evaluate complex cases using available evidence, context, and defined project requirements.
  • Classify, label, or annotate assigned data accurately and consistently.
  • Provide high-quality human evaluations to support AI system training, evaluation, and improvement.
  • Document decisions and supporting reasoning clearly when required.
  • Identify unclear or unusual cases and flag them according to defined project processes.
  • Apply detailed project guidelines consistently across assigned tasks.
  • Maintain high levels of quality, accuracy, consistency, and attention to detail.
Required Qualifications
  • Educational background or equivalent experience in Artificial Intelligence, Machine Learning, Computer Science, Data Science, Statistics, Engineering, or a related field.
  • Experience working with AI/ML systems, machine learning models, AI evaluation, data analysis, or technical system analysis.
  • Strong understanding of core AI and machine learning concepts.
  • Ability to analyse complex technical data, system behaviour, and patterns.
  • Experience evaluating AI systems, models, outputs, or data is preferred.
  • Familiarity with system telemetry, event data, logs, or other technical signals is preferred.
  • Strong analytical and problem-solving skills.
  • Ability to identify meaningful patterns and make informed decisions based on complex or incomplete information.
  • Strong attention to detail and the ability to apply detailed guidelines consistently.
  • Experience with data annotation, AI evaluation, human feedback, model evaluation, or structured data review is preferred.
  • Familiarity with bot detection, automated abuse detection, anomaly detection, fraud detection, or similar areas is preferred.
  • Strong written English comprehension and documentation skills.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Evaluator - Human-in-the-Loop (Freelance)
AI/ML Evaluator - Human-in-the-Loop (Freelance)

Welocalize • Town of Texas (WI)

On-site
USD 192,864,000 - 220,416,000
AI Evaluation Specialist
AI Evaluation Specialist

micro1 • United States

On-site
AUD 70,000 - 110,000
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Alabama

Remote
Flexible working hours
Competitive pay up to $80/hour
Experience in advanced AI projects
AI Agent Evaluation Analyst
AI Agent Evaluation Analyst

Mindrift • Dallas (TX)

Remote
Flexible remote work
Competitive pay up to $55/hour
Experience in advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Austin (TX)

Remote
Competitive pay
Flexible schedule
Experience in advanced AI projects
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Wisconsin

Remote
Competitive pay up to $60/hour
Flexible remote work
Opportunity to work on advanced AI projects
AI Agent Evaluation Analyst
AI Agent Evaluation Analyst

Mindrift • South Carolina

Remote
Flexible schedule
Experience in AI projects
Competitive pay
Machine Learning (ML) AI Task Auditor - Freelance AI Trainer Project
Machine Learning (ML) AI Task Auditor - Freelance AI Trainer Project

Triwill Group • Northern (KY)

Hybrid
USD 83,000 - 165,000
Remote work
Personalized AI Response Evaluation Rater
Personalized AI Response Evaluation Rater

OpenTrain AI • Northern (KY)

Hybrid
USD 17,000 - 28,000