A complete application in a minute — tailored resume and cover letter, ready to send.
OpenTrain AI, Inc. seeks a part-time contractor to create realistic terminal-based scientific tasks for AI training and evaluation. You will build reproducible benchmark environments and assess agents’ ability to reason, use tools, and debug calculations.
The role requires strong scientific programming expertise and independent validation of workflows, running in Linux/terminal environments with fixed dependencies.
You will create realistic terminal-based scientific tasks for AI training and evaluation. The work turns real physics, chemistry, materials science, astronomy, and computational science workflows into reproducible benchmark environments.
You will combine scientific modeling, software development, automated grading, and expert review. You will assess whether AI agents can reason through problems, use command-line tools, debug calculations, and produce reliable scientific files.
The listing does not specify a pay rate. This is a part-time contractor role with a 20+ hour weekly commitment, open worldwide.
Experience with scientific Python libraries, simulation methods, optimization, or statistical modeling is useful. Familiarity with research software engineering, benchmark creation, automated graders, AI coding-agent evaluation, publications, or open-source computational science also supports this work.
AI training work uses examples, tests, and human reviews to improve how AI systems behave. In this role, your scientific expertise helps evaluate whether AI agents can complete reliable computational workflows, which is why advanced subject knowledge and careful technical review matter.