Job Search and Career Advice Platform

Enable job alerts via email!

Evaluation Scenario Writer & AI Agent Testing Specialist

Mindrift

Remote

ZAR 200 000 - 300 000

Part time

Today
Be an early applicant

Generate a tailored resume in minutes

Land an interview and earn more. Learn more

Job summary

A tech staffing agency is looking for software engineers for project-based AI testing roles. Ideal candidates should have over 3 years of Python development experience and skills in creating structured test cases. This opportunity offers flexibility in work hours, with tasks expected to take 6-10 hours to complete. Pay rates vary, reaching up to $24/hour depending on expertise. English proficiency at B2 level is also required. Join a fast-paced environment focused on enhancing AI systems while working remotely.

Qualifications

  • 3+ years of software development experience, ideally focused on Python.
  • Experience with version control using Git.
  • Comfortable with structured formats like JSON/YAML.

Responsibilities

  • Create structured test cases simulating human workflows.
  • Define gold-standard behavior for evaluating actions.
  • Analyze agent logs and decision paths.

Skills

Software development
Python
Git
JSON/YAML
Understanding core LLM limitations
Docker
Job description

Please submit your CV in English and indicate your level of English proficiency.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

What This Opportunity Involves

While each project involves unique tasks, contributors may:

  • Create structured test cases that simulate complex human workflows
  • Define gold-standard behavior and scoring logic to evaluate actions
  • Analyze agent logs, failure modes, and decision paths
  • Work with code repositories and test frameworks to validate your scenarios
  • Iterate on prompts, instructions, and test cases to improve clarity and difficulty
  • Ensure that scenarios are production-ready, easy to run, and reusable
What We Look For

This opportunity is a good fit for software engineers, open to part-time, non-permanent projects. Ideally, contributors will have:

  • 3+ years of software development experience with strong Python focus
  • Experience with Git and code repositories
  • Comfortable with structured formats like JSON/YAML for scenario description
  • Understanding core LLM limitations (hallucinations, bias, context limits) and how these affect evaluation design
  • Familiarity with Docker
  • English proficiency - B2
How It Works

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

Project time expectations

Tasks for this project are estimated to take 6-10 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Payment
  • Paid contributions, with rates up to $24/hour*
  • Fixed project rate or individual rates, depending on the project
  • Some projects include incentive payments
  • Note: Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be offered to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project
Get your free, confidential resume review.
or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.