An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Planet Pharma seeks experienced scientists and engineers to evaluate how frontier AI models handle real technical work. You design challenging tasks, analyze data, and judge model outputs against professional standards while assembling test plans, data analyses, and technical reports.
You will run tasks through frontier AI agents, grade results with detailed rubrics, and identify concrete failures with evidence. Strong coding and data analysis skills are essential for this fully remote role.
is looking for experienced scientists and engineers to evaluate how frontier AI models handle real technical work: analyzing test, measurement or process data, sizing and verifying a design, designing an experiment and reading it out, interpreting a simulation, writing the technical report. You bring the judgment you have built catching the unit error in a test file, the multiple comparisons trap in a process improvement study, and the assumption that does not hold at the boundary. We bring the model output that judgment is needed to grade.
is looking for experienced scientists and engineers to evaluate how frontier AI models handle real technical work: analyzing test, measurement or process data, sizing and verifying a design, designing an experiment and reading it out, interpreting a simulation, writing the technical report. You bring the judgment you have built catching the unit error in a test file, the multiple comparisons trap in a process improvement study, and the assumption that does not hold at the boundary. We bring the model output that judgment is needed to grade.
In this role, you will design challenging, realistic tasks drawn from your own practice, such as a calculation package with stated assumptions and checks, a test plan and data analysis, a failure or deviation investigation, a design trade study, an experimental protocol with acceptance criteria, a circuit or system design review, or a technical report, run them through frontier AI agents, and evaluate what comes back against a professional standard.
Across our STEM and Data Science programs, tasks are grounded in real day-to-day workflows and checked against frontier models so only genuinely hard tasks make it through. Some projects are authored and verified in code, so coding / scientific computing is a strong plus and is required on those projects.
You will work with realistic professional files, the kind a practitioner in your field actually handles, which you assemble yourself. Some tasks are compact, built around a handful of files; others are larger scenarios that take several days to build. In every case the goal is the same: a task a competent professional in your field would complete correctly and a frontier model currently gets wrong.
This is not a traditional science or engineering role. You will be helping build better AI by putting your knowledge to work in a structured, flexible, fully remote environment. The work is long-form and self-directed, and clear written reasoning matters as much as technical depth.