Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Mercor in the United Kingdom is seeking PhD and Master's physicists to author AI evaluation tasks for a new frontier benchmark. You will create original, executable problems that current models struggle to solve, focusing on at least two physics subdomains and leveraging Python or R for computation.
This six-week, part-time engagement (20+ hours/week) starts immediately, with opportunities to publish and contribute to a growing AI benchmarking effort.
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)
Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.
Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
Write scientific prompts based on the input
Build the grading criteria that define a correct answer
Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed
PhD in physics, applied physics, or a closely related field
Demonstrated depth in at least two of the following subdomains: condensed matter, optics, quantum information/computing, computational physics, astrophysics, particle physics
Working proficiency in Python, R, or another relevant programming language for scientific computing
Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks
Publications in peer-reviewed journals
Prior scientific software or research engineering experience
Duration: 6 weeks
Commitment: part-time, 20+ hours per week
Start date: immediate
Upload your resume and application form
A 25-minute conversational interview covering your background, experience, and motivations
Follow up within a few days with next steps and onboarding