Senior AI Agent Evaluation Engineer

aitrainer

Deutschland

Vor Ort

EUR 34.453 - 60.293

Teilzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

aitrainer is seeking experienced software developers for project-based AI opportunities aimed at improving AI systems. This role involves building realistic developer environments, designing and refining tasks, and writing effective tests for agent solutions. Candidates should have over 5 years of experience in software development, particularly with Python and JavaScript, and possess a strong proficiency in English.

The opportunity allows flexibility in scheduling, with compensation up to $50/hour depending on experience and project pace.

Qualifikationen

  • 5+ years in software development.
  • Experience writing functional and integration tests.
  • Fluent in English (B2+).

Aufgaben

  • Build realistic developer environments with codebase and infrastructure.
  • Design tasks from intermediate states of environments.
  • Write tests that verify agent solutions effectively.
  • Iterate on tasks based on QA feedback.

Kenntnisse

Software development
Python (FastAPI)
JavaScript/TypeScript (React)
Docker
Postgres
Kafka
Redis
Test writing (functional, integration)
English proficiency (B2+)

Jobbeschreibung

Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

What this opportunity involves
  • Build realistic developer environments – a virtual company with codebase, infrastructure, and context that forms a believable development history.
  • Design tasks from intermediate states of these environments – craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent.
  • Write tests that verify agent solutions – accept all valid approaches and reject incorrect ones, neither too strict nor too lenient.
  • Iterate on tasks and tests based on QA feedback – review agent solutions, analyze failures, and refine until the evaluation is fair and robust.
What this is NOT
  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch – the agent writes most of the code; you guide and evaluate.
What we look for
  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency – B2+
Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non‑trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions – writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How it works
  • Apply
  • Pass qualification(s)
  • Join a project
  • Complete tasks
  • Get paid
Effort estimate

Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Compensation

Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Software Engineer - Agent Evaluation
Senior Software Engineer - Agent Evaluation

aitrainer • Deutschland

Hybrid
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Berlin

Remote
Competitive hourly rates
Flexible schedule
Experience in advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Deutschland

Remote
Competitive pay up to $52/hour
Flexible schedule
Experience with advanced AI projects
python developer for AI coding agents
python developer for AI coding agents

Enfint • Deutschland

Remote
EUR 108.000 - 179.000
Гибкий график
Проектная занятость
Оплата до $150/ч
+1
Freelance AI Red Team Engineer
Freelance AI Red Team Engineer

Mindrift • Berlin

Remote
Competitive hourly rates
Flexible working hours
Experience with advanced AI projects
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • Köln

Vor Ort
Freelance project-based collaboration
Flexible working hours
Competitive compensation
Freelance Software Developer (Ruby) - AI Trainer
Freelance Software Developer (Ruby) - AI Trainer

Mindrift • Berlin

Remote
Flexible scheduling
Competitive hourly rates up to $52
Opportunity to influence AI development
+1
Freelance Financial Analyst - AI Trainer
Freelance Financial Analyst - AI Trainer

Mindrift • Köln

Remote
Flexible working hours
Competitive pay based on expertise
Remote work on advanced AI projects
Senior Portfolio Manager - Freelance AI Trainer
Senior Portfolio Manager - Freelance AI Trainer

aitrainer • Deutschland

Hybrid
Freelance Financial Analyst - AI Trainer
Freelance Financial Analyst - AI Trainer

Mindrift • Hamburg

Remote
Flexible working hours
Competitive pay up to $50/hour
Experience with advanced AI projects