Get more replies from employers
Send a job-specific resume in minutes.
Mercor is recruiting PhD and Master's scientists to author AI evaluation tasks for Sci Code, a benchmark project in scientific computing. You will craft original, executable research problems that current frontier models cannot solve, with a focus on at least two deep subdomains and a coding emphasis.
Duration is 6 weeks, part-time at 20+ hours per week, with immediate start. Candidates should have strong Python for scientific computing, experience with GitHub and Docker workflows, and
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)
Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.
Domains - depth required in at least two subdomains (with a coding focus)
What you'll do
Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
Write scientific prompts based on the input
Build the grading criteria that define a correct answer
Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed
Required
PhD in mathematics, applied mathematics, computational mathematics, or a closely related field
Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance
Working proficiency in Python for scientific computing
Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks
Preferred
Publications in peer-reviewed journals
Prior scientific software or research engineering experience
Engagement
Duration: 6 weeks
Commitment: part-time, 20+ hours per week
Start date: immediate