Remote AI Task Engineer for Real-World Agent Scenarios

S27a

United States

Remote

USD 15,000 - 35,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Flexible hours
Fully remote
Paid pilot task

Job summary

Careerflow Human Data Labs seeks a detail-oriented contractor to design realistic, multi-step scenarios for AI evaluation. You will assemble documents and data, provide precise instructions to AI agents, and write Python scripts to automate environment setup and scoring. Strong Linux/VM experience and Git basics are essential.

Work is fully remote with a 3–4 week engagement, starting immediately. The role offers flexible hours and a quick onboarding process, with paid pilot tasks and ongoing

Qualifications

  • Proficient in writing Python scripts to set up test environments and scoring pipelines.
  • Experience using AI tools and models to develop evaluation tasks.
  • Familiarity with AI/ML projects and agent-based benchmarks preferred.
  • Comfortable working on Linux desktops and hosted VMs.
  • Strong attention to detail to ensure clear, verifiable tasks.

Responsibilities

  • Design realistic, multi-step scenarios across several applications.
  • Assemble necessary files: spreadsheets, documents, emails, data.
  • Write clear instructions for AI agents with no giveaways.
  • Develop Python scripts to automate setup and scoring of tasks.
  • Test tasks, run experiments, and document the process.

Skills

Python
Git
OS/Linux terminal
Attention to detail
AI tooling familiarity

Tools

Ubuntu
Virtual machines

Job description

Careerflow Human Data Labs seeks a detail-oriented contractor to design realistic, multi-step scenarios for AI evaluation. You will assemble documents and data, provide precise instructions to AI agents, and write Python scripts to automate environment setup and scoring. Strong Linux/VM experience and Git basics are essential.

Work is fully remote with a 3–4 week engagement, starting immediately. The role offers flexible hours and a quick onboarding process, with paid pilot tasks and ongoing

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Evaluation Scenario Designer
Remote AI Evaluation Scenario Designer

Mindrift • United States

Remote
USD 40,000 - 80,000
Remote: AI Agent Evaluation Scenario Designer
Remote: AI Agent Evaluation Scenario Designer

Mindrift • Kentucky

Remote
USD 10,000 - 60,000
Remote AI Training Architect for Software Tasks
Remote AI Training Architect for Software Tasks

YO AI Labs • New York (NY)

Remote
USD 35,000 - 65,000
Remote RL Scenario Architect
Remote RL Scenario Architect

Terac • United States

On-site
USD 106,000 - 143,000
Senior Software Engineer (Remote) – AI Task Environments
Senior Software Engineer (Remote) – AI Task Environments

YO AI Labs • Los Angeles (CA)

Remote
USD 83,000 - 165,000
Evaluation Scenario Writer - AI Agent Testing Specialist
Evaluation Scenario Writer - AI Agent Testing Specialist

Mindrift • Houston (TX)

Remote
Get paid for your expertise
Flexible work schedule
Opportunity to work on advanced AI projects
Evaluation Scenario Writer - AI Agent Testing Specialist
Evaluation Scenario Writer - AI Agent Testing Specialist

Mindrift • Kentucky

Remote
Flexible schedule
Competitive hourly pay up to $60
Remote work opportunity
+1
Remote Dev & Infra Engineer, AI Workflow Evaluator
Remote Dev & Infra Engineer, AI Workflow Evaluator

YO AI Labs • Illinois

Remote
USD 110,000 - 207,000
AI Agent Evaluation Engineer - Remote
AI Agent Evaluation Engineer - Remote

YO AI Labs • Boston (MA)

Remote
USD 42,000 - 65,000
Remote Senior AI Agent Evaluation Engineer
Remote Senior AI Agent Evaluation Engineer

YO AI Labs • Boston (MA)

Remote
USD 42,000 - 65,000