Remote AI Evaluation Architect - Tech Docs & Code

Weekday 1

United States

Remote

USD 21,000 - 28,000

Part time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible hours
Weekly payments

Job summary

Weekday 1 is seeking a seasoned technology professional to design high-quality benchmark tasks for evaluating AI systems on software engineering and data science workflows. This fully remote, independent contractor role focuses on creating realistic, multi-step tasks grounded in technical documentation, codebases, APIs, and architectures.

You will develop evaluation standards, ground-truth solutions, and scoring rubrics, collaborating with research teams while maintaining rigorous accuracy and

Qualifications

  • Minimum 3 years of hands-on professional experience in software engineering, data science, or data analytics.
  • Strong understanding of technical documentation, software development workflows, and engineering best practices.
  • Experience with codebases, APIs, technical specifications, or system architecture documentation.
  • Excellent analytical thinking and problem-solving skills.
  • Strong written communication with the ability to create clear technical instructions and evaluation criteria.
  • Ability to work independently with high accuracy and consistency.

Responsibilities

  • Design AI evaluation tasks based on professional technology workflows.
  • Develop ground-truth solutions and scoring rubrics for tasks.
  • Contribute domain expertise from software engineering, data science, or analytics.
  • Collaborate with research teams to improve benchmark quality.
  • Refine tasks based on feedback and evolving requirements.

Skills

Software Engineering
Data Science
Data Analytics

Job description

Weekday 1 is seeking a seasoned technology professional to design high-quality benchmark tasks for evaluating AI systems on software engineering and data science workflows. This fully remote, independent contractor role focuses on creating realistic, multi-step tasks grounded in technical documentation, codebases, APIs, and architectures.

You will develop evaluation standards, ground-truth solutions, and scoring rubrics, collaborating with research teams while maintaining rigorous accuracy and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Evaluation Specialist
Remote AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
AI Evaluation Architect - Data Science Expert (Remote)
AI Evaluation Architect - Data Science Expert (Remote)

Weekday AI (YC W21) • United States

On-site
USD 165,312 - 234,192
Fully remote
Weekly payments
Technology AI Evaluation Expert
Technology AI Evaluation Expert

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Remote AI Benchmark Technical Writer
Remote AI Benchmark Technical Writer

YO AI Labs • Maryland

Remote
USD 55,000 - 96,000
Remote AI Benchmark Technical Writer (Contractor)
Remote AI Benchmark Technical Writer (Contractor)

YO AI Labs • California (MO)

Remote
USD 55,000 - 110,000
Remote Technical Writer for AI Benchmark Tasks
Remote Technical Writer for AI Benchmark Tasks

YO AI Labs • Town of Texas (WI)

Remote
USD 60,000 - 90,000
Remote AI Legal Benchmark Architect
Remote AI Legal Benchmark Architect

Weekday 1 • United States

Remote
USD 114,000 - 145,000
Remote Technical Writer for AI Benchmark & Documentation
Remote Technical Writer for AI Benchmark & Documentation

YO AI Labs • San Francisco (CA)

Remote
USD 55,000 - 117,000
Remote AI Consulting Quality Evaluator
Remote AI Consulting Quality Evaluator

Weekday 1 • United States

Remote
USD 207,000 - 303,000