AI Benchmark Architect (Remote) | Python

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Fully remote
Weekly payments

Job summary

Weekday 1 is seeking experienced Software Engineers for a fully remote, contractor engagement to build multi-step software engineering benchmarks for frontier AI models. You will implement, debug, configure environments, and reason about tasks alongside AI researchers.

The role emphasizes Python reference solutions, reproducibility, and rigorous validation checks, with approximately 35 hours per week. Weekly payments apply for this remote engagement.

Qualifications

  • Master's degree, PhD, or equivalent practical experience in CS/Software Engineering or STEM
  • Minimum 1 year of professional software engineering or research engineering experience
  • Strong Python proficiency including design, scripting, debugging, testing, and clean code
  • Solid experience with Git, IDEs, and modern software development workflows
  • Good understanding of software architecture, debugging methods, and best practices
  • Experience with AI coding assistants, prompt engineering, or AI agent workflows is preferred

Responsibilities

  • Design realistic, multi-step software engineering benchmarks reflecting typical development workflows
  • Develop comprehensive Python reference solutions that are reproducible and validated
  • Use AI coding assistants as part of the workflow while evaluating their strengths and failure modes
  • Review benchmark tasks for technical correctness, clarity, completeness, and difficulty
  • Analyze AI-generated solutions to identify errors, debugging challenges, and reasoning gaps
  • Collaborate with AI researchers and engineering teams to refine evaluation methodologies

Skills

Python
Git
Debugging
Software design
Code testing

Education

Master's degree / PhD in CS or STEM

Tools

IDEs
AI coding assistants

Job description

Weekday 1 is seeking experienced Software Engineers for a fully remote, contractor engagement to build multi-step software engineering benchmarks for frontier AI models. You will implement, debug, configure environments, and reason about tasks alongside AI researchers.

The role emphasizes Python reference solutions, reproducibility, and rigorous validation checks, with approximately 35 hours per week. Weekly payments apply for this remote engagement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Data Science & AI Benchmark Analyst
Remote Data Science & AI Benchmark Analyst

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
Remote QA/Test Engineer — AI Benchmark Validation
Remote QA/Test Engineer — AI Benchmark Validation

Weekday AI (YC W21) • United States

On-site
USD 83,000 - 124,000
Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
Remote STEM Researcher - AI Evaluation & Benchmark Design
Remote STEM Researcher - AI Evaluation & Benchmark Design

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Benchmarking Software Engineer (Remote)
Benchmarking Software Engineer (Remote)

Office Hours • United States

Hybrid
USD 160,000 - 210,000
Competitive salary and equity
Medical, dental, and vision coverage
401(k)
+4
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
AI Benchmark Auditor (Python & Evaluation)
AI Benchmark Auditor (Python & Evaluation)

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote AI Evaluation Architect - Tech Docs & Code
Remote AI Evaluation Architect - Tech Docs & Code

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
Senior Software Engineer - Remote AI Training Environments
Senior Software Engineer - Remote AI Training Environments

Weekday 1 • United States

Remote
USD 138,000 - 317,000
AI Evaluation Architect - Data Science Expert (Remote)
AI Evaluation Architect - Data Science Expert (Remote)

Weekday AI (YC W21) • United States

On-site
USD 165,312 - 234,192
Fully remote
Weekly payments