SWE-Bench AI Task Auditor - Freelance AI Trainer Project

Jobgether SRL

United States

Remote

USD 68,000 - 97,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote freelance contract
Flexible project-based work
Impact on AI training workflows

Job summary

Jobgether SRL, on behalf of a partner, seeks a SWE-Bench AI Task Auditor - Freelance AI Trainer Project in the United States. Fully remote, project-based, you review SWE tasks for accuracy, realism, solvability, and reproducibility.

You will troubleshoot codebase integrations, test failures, logic issues, and provide clear feedback to task creators to improve AI training workflows.

Qualifications

  • Professional software engineers with real-world app experience.
  • Strong grasp of software engineering principles and workflows.
  • Ability to navigate complex codebases and understand integrations.

Responsibilities

  • Evaluate SWE tasks for technical accuracy, realism, solvability, reproducibility, and alignment with SWE-Bench standards.
  • Review task codebases, integrations, tests, and evaluation criteria to identify weaknesses.
  • Rigorously test and troubleshoot complex technical scenarios to confirm expected behavior.
  • Investigate codebase integration problems, test failures, logic errors, and implementation issues.
  • Provide clear, actionable feedback enabling task creators to fix problems.
  • Apply professional software engineering judgment to assess realism of development scenarios.
  • Help maintain high technical rigor across AI training and evaluation workflows.

Skills

Software engineering
Codebase navigation
Analytical thinking
Troubleshooting
Clear communication

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SWE-Bench AI Task Auditor - Freelance AI Trainer Project based in United States.

This freelance opportunity is designed for experienced software engineers who want to contribute their technical expertise to the development and evaluation of AI systems.

You will review software engineering tasks used to train and assess advanced AI models.

Your work will focus on ensuring tasks are technically accurate, realistic, reproducible, and appropriately tested.

You will investigate codebase integrations, test failures, logic issues, and other technical challenges.

The role offers flexibility through a fully remote, project-based working model.

You will apply practical software engineering judgment rather than simply following predefined checks.

Your feedback will directly help improve the quality and reliability of AI training workflows.

Accountabilities
  • Evaluate software engineering tasks for technical accuracy, realism, solvability, reproducibility, and alignment with SWE-Bench standards.

  • Review task codebases, integrations, tests, and evaluation criteria to identify potential technical weaknesses.

  • Rigorously test and troubleshoot complex technical scenarios to determine whether tasks function as intended.

  • Investigate codebase integration problems, test failures, logic errors, and other implementation issues.

  • Provide clear, precise, and actionable feedback that enables task creators to correct identified problems.

  • Apply professional software engineering judgment to assess whether tasks reflect realistic development scenarios.

  • Help maintain a high standard of technical rigor and accuracy across AI training and evaluation workflows.

Requirements
  • Demonstrable professional experience in software engineering, including experience navigating complex codebases and developing real-world applications.

  • Strong knowledge of software engineering principles and familiarity with SWE-Bench-style tasks and evaluation workflows.

  • Strong analytical and problem-solving abilities, with the capacity to investigate complex technical scenarios systematically.

  • Experience troubleshooting code, diagnosing test failures, and identifying underlying logic or integration issues.

  • Ability to evaluate technical work objectively and communicate findings clearly and constructively.

  • Strong attention to detail and a rigorous approach to technical validation.

  • Ability to work independently and manage project-based assignments effectively.

  • Deep expertise in one relevant software engineering specialty is sufficient; expertise across multiple domains is not required.

  • A secure computer and reliable, high-speed internet connection suitable for remote technical work.

Benefits
  • Fully remote freelance contract.

  • Flexible project-based working environment.

  • Opportunity to contribute directly to the training and evaluation of AI systems.

  • Ability to apply real-world software engineering expertise to challenging technical tasks.

  • Compensation of $60/hour, with the final rate determined based on experience, expertise, and geographic location.

  • No company-sponsored health insurance, PTO, or other employee benefits, as this is a freelance contractor position.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SWE-Bench AI Task Auditor - Freelance AI Trainer Project (Fully Remote)
SWE-Bench AI Task Auditor - Freelance AI Trainer Project (Fully Remote)

Jobgether SRL • United States

Remote
USD 100,000 - 140,000
Remote SWE-Bench AI Task Auditor - Freelance
Remote SWE-Bench AI Task Auditor - Freelance

Meridial • Austin (TX)

On-site
USD 68,000 - 97,000
Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
SWE-Bench Task Auditor · Mercor Mercor · $70-90/hr · remote in US · 3w ago $70-90/hr remote in US 3w ago
SWE-Bench Task Auditor · Mercor Mercor · $70-90/hr · remote in US · 3w ago $70-90/hr remote in US 3w ago

Benture • Northern (KY)

Hybrid
USD 96,000 - 124,000
Backend Software Engineer - AI Trainer
Backend Software Engineer - AI Trainer

DataAnnotation • Colorado

Remote
USD 69,000 - 138,000
Project selection
Cutting-edge AI coding agents
Remote work from home
+1
Backend Software Engineer - AI Trainer
Backend Software Engineer - AI Trainer

DataAnnotation • Missouri

Remote
USD 69,000 - 138,000
Choose projects and schedule
Work with unreleased AI agents
Work from home
+1
Senior SWE - AI Specialist
Senior SWE - AI Specialist

Mercor • Lisbon (AR)

Remote
USD 207,000 - 289,000
Paid weekly via Stripe Connect
Senior Software Engineer – Open Source & SWE-Bench Evaluation (Fully Remote)
Senior Software Engineer – Open Source & SWE-Bench Evaluation (Fully Remote)

Anyone AI Inc. • United States

Remote
USD 83,000 - 165,000
Technology AI Evaluation Expert
Technology AI Evaluation Expert

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour

24-MAG • United States

Remote
USD 69,000 - 138,000
Remote work
Flexible hours
Contractor engagement