An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Weekday 1 is seeking researchers to develop rigorous evaluation benchmarks for frontier AI models. The role focuses on transforming experimental design and data-driven analysis into sophisticated, multi-step benchmarks that challenge current AI systems.
You will work with AI researchers to uncover subtle reasoning errors and methodological flaws while contributing to the continuous refinement of evaluation methodologies. This is a fully remote, full-time engagement with about 35 hours per week.
Weekday 1 is seeking researchers to develop rigorous evaluation benchmarks for frontier AI models. The role focuses on transforming experimental design and data-driven analysis into sophisticated, multi-step benchmarks that challenge current AI systems.
You will work with AI researchers to uncover subtle reasoning errors and methodological flaws while contributing to the continuous refinement of evaluation methodologies. This is a fully remote, full-time engagement with about 35 hours per week.