Technology AI Evaluation Expert

Weekday AI

United States

Remote

USD 83,000 - 103,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Weekday AI is seeking an experienced contractor for an advanced AI research initiative focused on how professionals’ documents are understood and how instructions are executed. You will design realistic, multi-step benchmark tasks based on technical documentation, code repositories, API references, and architecture diagrams to evaluate AI performance in software engineering and data science.

This fully remote, independent role offers flexible hours (15–20 hours per week) and weekly payments.

Qualifications

  • Minimum 3 years of hands-on professional experience in software engineering, data science, or data analytics.
  • Experience with codebases, APIs, technical specifications, or system architecture documentation.
  • Strong written communication and ability to create clear technical instructions and evaluation criteria.
  • Ability to work independently with high accuracy and consistency.

Responsibilities

  • Design AI evaluation tasks based on professional technology workflows.
  • Develop challenges using technical specifications, architecture documents, API documentation, codebases, web research, and code execution.
  • Ensure each task includes a clearly defined output and evaluation criteria.
  • Write ground-truth solutions and structured scoring rubrics.
  • Collaborate with research teams to improve benchmark quality and consistency.
  • Operate as an independent contractor with flexible remote hours (15–20 hours per week).

Skills

Software Engineering
Data Science
Data Analytics

Job description

This role is for one of our clients

Compensation: $60-$75 per hour

Join an advanced AI research initiative focused on improving how next-generation AI systems understand professional documents, execute complex instructions, and reason through real-world technical workflows. We are seeking experienced technology professionals to design high-quality benchmark tasks that evaluate AI performance across software engineering and data science domains.

In this role, you will create realistic, multi-step evaluation tasks based on technical documentation, code repositories, API references, architecture diagrams, and other workplace resources. Your work will help measure and improve the ability of AI models to interpret technical information, follow detailed instructions, and generate accurate, well-structured outputs.

This is a fully remote, independent contractor opportunity with flexible working hours.

Key Responsibilities
Design AI Evaluation Tasks
  • Create realistic, multi-step benchmark tasks based on professional technology workflows.
  • Develop challenges using technical specifications, architecture documents, API documentation, codebases, web research, and code execution.
  • Ensure each task includes a clearly defined expected output and objective evaluation criteria.
Develop Evaluation Standards
  • Write comprehensive ground-truth solutions and structured scoring rubrics.
  • Design tasks that assess reasoning, technical understanding, instruction following, and output quality.
  • Maintain high standards of technical accuracy, clarity, and reproducibility.
Contribute Domain Expertise
  • Apply real-world knowledge from software engineering, data science, or analytics to create authentic evaluation scenarios.
  • Collaborate with research teams to improve benchmark quality and consistency.
  • Continuously refine tasks based on project feedback and evolving evaluation requirements.
Required Qualifications
  • Minimum 3 years of hands-on professional experience in one or more of the following areas:
    • Software Engineering
    • Data Science
    • Data Analytics
  • Strong understanding of technical documentation, software development workflows, and engineering best practices.
  • Experience working with codebases, APIs, technical specifications, or system architecture documentation.
  • Excellent analytical thinking and problem-solving skills.
  • Strong written communication with the ability to create clear technical instructions and evaluation criteria.
  • Ability to work independently while maintaining high standards of accuracy and consistency.
Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of 15–20 hours per week.
  • Projects may be extended, shortened, or concluded based on business needs and performance.
  • Weekly payments processed through supported payment platforms.
Why Join
  • Help shape the next generation of AI systems for technical reasoning and document understanding.
  • Work on intellectually challenging projects involving real-world engineering and data science workflows.
  • Apply your technical expertise to improve advanced AI evaluation benchmarks.
  • Enjoy flexible remote work with meaningful impact on AI research.
Equal Opportunity Statement

We are committed to providing equal opportunities to all qualified applicants without regard to legally protected characteristics. Reasonable accommodations are available upon request.

Contract Information
  • Independent contractor engagement.
  • Fully remote work completed on your own schedule.
  • Weekly payments are processed based on approved work completed.
  • Work does not involve access to confidential or proprietary information from any employer, client, or institution.
  • Please note that visa sponsorship is not available for this opportunity.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Specialist
AI Evaluation Specialist

Weekday AI • United States

Remote
USD 83,000 - 110,000
QA/Test Engineer
QA/Test Engineer

Weekday AI • United States

Remote
USD 83,000 - 124,000
Remote AI Evaluation Architect - Tech Docs & Code
Remote AI Evaluation Architect - Tech Docs & Code

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Generalist Expert (UK/Europe)
Generalist Expert (UK/Europe)

Weekday AI • United States

Remote
USD 69,000 - 96,000
Fully remote
Data Science Expert
Data Science Expert

Weekday AI • United States

Remote
USD 165,000 - 234,000
Fully remote
Weekly payments
Software Engineering Expert
Software Engineering Expert

Weekday AI • United States

Remote
USD 83,000 - 124,000
Legal AI Evaluation Expert
Legal AI Evaluation Expert

Weekday AI • United States

Remote
USD 138,000 - 207,000
Remote work
Flexible schedule
Weekly payments
Remote | QA Engineer — $90–$175/hour
Remote | QA Engineer — $90–$175/hour

engineeringjobs.net, Inc. • United States

Remote
USD 124,000 - 241,000
Data Analyst (AI Evaluation)
Data Analyst (AI Evaluation)

Jobgether SRL • United States

Remote
USD 65,000 - 90,000
Flexible remote work
Competitive compensation
Contract-based project work
+1
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday AI • United States

Remote
USD 83,000 - 124,000