Erhalte mehr Antworten von Arbeitgebern
Versende in nur wenigen Minuten einen passgenauen Lebenslauf.
Aitrainer is looking for a candidate to author complex tasks for AI model evaluation within the Technology domain. The role requires at least 15–20 hours of commitment per week.
The ideal candidate has over 3 years of experience in software engineering or data science, with a focus on creating precise evaluation tasks based on technical documentation. Responsibilities include engaging with complex requests and enabling accurate AI outputs.
We are building a benchmark dataset to evaluate AI models on professional document understanding and instruction following within the Technology domain.
Tasks consist of complex, multi-step requests grounded in real-world workspace files (technical specs, architecture docs, API references, codebases), web search, and code execution — each paired with a clearly defined ground truth output and an objective evaluation rubric. You will be responsible for authoring tasks that test an AI's ability to reason over technical documentation, follow precise instructions, and produce accurate, well-structured outputs.
We expect a minimum commitment of 15–20 hours per week.
Ideal candidates have 3+ years of hands‑on experience in one or more of the following sub-domains:
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.