Technology AI Evaluation Expert

Weekday 1

United States

Remote

USD 21,000 - 28,000

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible hours
Weekly payments

Job summary

Weekday 1 is seeking a seasoned technology professional to design high-quality benchmark tasks for evaluating AI systems on software engineering and data science workflows. This fully remote, independent contractor role focuses on creating realistic, multi-step tasks grounded in technical documentation, codebases, APIs, and architectures.

You will develop evaluation standards, ground-truth solutions, and scoring rubrics, collaborating with research teams while maintaining rigorous accuracy and

Qualifications

  • Minimum 3 years of hands-on professional experience in software engineering, data science, or data analytics.
  • Strong understanding of technical documentation, software development workflows, and engineering best practices.
  • Experience with codebases, APIs, technical specifications, or system architecture documentation.
  • Excellent analytical thinking and problem-solving skills.
  • Strong written communication with the ability to create clear technical instructions and evaluation criteria.
  • Ability to work independently with high accuracy and consistency.

Responsibilities

  • Design AI evaluation tasks based on professional technology workflows.
  • Develop ground-truth solutions and scoring rubrics for tasks.
  • Contribute domain expertise from software engineering, data science, or analytics.
  • Collaborate with research teams to improve benchmark quality.
  • Refine tasks based on feedback and evolving requirements.

Skills

Software Engineering
Data Science
Data Analytics

Job description

This role is for one of our clients

Compensation: $60-$75 per hour

Join an advanced AI research initiative focused on improving how next-generation AI systems understand professional documents, execute complex instructions, and reason through real-world technical workflows. We are seeking experienced technology professionals to design high-quality benchmark tasks that evaluate AI performance across software engineering and data science domains.

In this role, you will create realistic, multi-step evaluation tasks based on technical documentation, code repositories, API references, architecture diagrams, and other workplace resources. Your work will help measure and improve the ability of AI models to interpret technical information, follow detailed instructions, and generate accurate, well-structured outputs.

This is a fully remote, independent contractor opportunity with flexible working hours.

Requirements

Key Responsibilities
Design AI Evaluation Tasks
  • Create realistic, multi-step benchmark tasks based on professional technology workflows.
  • Develop challenges using technical specifications, architecture documents, API documentation, codebases, web research, and code execution.
  • Ensure each task includes a clearly defined expected output and objective evaluation criteria.
Develop Evaluation Standards
  • Write comprehensive ground-truth solutions and structured scoring rubrics.
  • Design tasks that assess reasoning, technical understanding, instruction following, and output quality.
  • Maintain high standards of technical accuracy, clarity, and reproducibility.
Contribute Domain Expertise
  • Apply real-world knowledge from software engineering, data science, or analytics to create authentic evaluation scenarios.
  • Collaborate with research teams to improve benchmark quality and consistency.
  • Continuously refine tasks based on project feedback and evolving evaluation requirements.
Required Qualifications
  • Minimum 3 years of hands-on professional experience in one or more of the following areas:
    • Software Engineering
    • Data Science
    • Data Analytics
  • Strong understanding of technical documentation, software development workflows, and engineering best practices.
  • Experience working with codebases, APIs, technical specifications, or system architecture documentation.
  • Excellent analytical thinking and problem-solving skills.
  • Strong written communication with the ability to create clear technical instructions and evaluation criteria.
  • Ability to work independently while maintaining high standards of accuracy and consistency.
Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of 15–20 hours per week.
  • Projects may be extended, shortened, or concluded based on business needs and performance.
  • Weekly payments processed through supported payment platforms.
Why Join
  • Help shape the next generation of AI systems for technical reasoning and document understanding.
  • Work on intellectually challenging projects involving real-world engineering and data science workflows.
  • Apply your technical expertise to improve advanced AI evaluation benchmarks.
  • Enjoy flexible remote work with meaningful impact on AI research.
Equal Opportunity Statement

We are committed to providing equal opportunities to all qualified applicants without regard to legally protected characteristics. Reasonable accommodations are available upon request.

Contract Information
  • Independent contractor engagement.
  • Fully remote work completed on your own schedule.
  • Weekly payments are processed based on approved work completed.
  • Work does not involve access to confidential or proprietary information from any employer, client, or institution.
  • Please note that visa sponsorship is not available for this opportunity.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Specialist
AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
Remote | AI Software Engineering Domain Expert — $100–$200/hour
Remote | AI Software Engineering Domain Expert — $100–$200/hour

24-Mag Llc • Northern (KY), New York (NY)

Hybrid
USD 138,000 - 276,000
AI Evaluation Expert - Remote
AI Evaluation Expert - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 76,000
Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Data Science Expert
Data Science Expert

Weekday 1 • United States

Remote
USD 165,000 - 234,000
Fully remote
Flexible scheduling
AI Evaluation Specialist - Remote
AI Evaluation Specialist - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 55,000
Remote AI Evaluation Architect - Tech Docs & Code
Remote AI Evaluation Architect - Tech Docs & Code

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Remote | QA Engineer — $90–$175/hour
Remote | QA Engineer — $90–$175/hour

engineeringjobs.net, Inc. • United States

Remote
USD 124,000 - 241,000
Remote | Technical Writer / Editor — $90–$140/hour
Remote | Technical Writer / Editor — $90–$140/hour

24-Mag Llc • Northern (KY)

Hybrid
USD 124,000 - 193,000
Remote contractor
Open to US, Canada, UK
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments