Stand out for this role — generate a tailored resume and cover letter in about a minute.
Planet Pharma seeks experienced legal professionals to evaluate how frontier AI models handle real legal work, including memos, contract review, drafting, and compliance analyses. You will design tasks, run them through frontier AI agents, and assess outputs against professional standards in a fully remote setting.
You will work with realistic practitioner files, building scenarios that test the AI’s ability to meet legal standards across litigation, corporate, compliance, and IP work.
is looking for experienced legal professionals to evaluate how frontier AI models handle real legal work: research memos, contract review and drafting, motion and brief sections, compliance analyses, due diligence summaries and regulatory responses. You bring the judgment you have built marking up a junior’s memo, checking a research trail before relying on it, and redlining a draft until it is the one you would sign. We bring the model output that judgment is needed to grade.
In this role, you will design challenging, realistic tasks drawn from your own practice, such as a research memo on a discrete question, a contract redline with an issues list, a section of a motion or brief, a compliance gap analysis, a due diligence summary, a discovery plan, or a response to a regulator, run them through frontier AI agents, and evaluate what comes back against a professional standard.
You will work with realistic professional files, the kind a practitioner in your field actually handles, which you assemble yourself. Some tasks are compact, built around a handful of files; others are larger scenarios that take several days to build. In every case the goal is the same: a task a competent professional in your field would complete correctly and a frontier model currently gets wrong.
This is not a traditional law role. You will be helping build better AI by putting your knowledge to work in a structured, flexible, fully remote environment. The work is long form and self directed, and clear written reasoning matters as much as technical depth.
Design challenging, realistic legal tasks drawn from your own day to day work: the scenario, a prompt phrased the way you would brief a trusted colleague, and the supporting files a professional would need (agreements, pleadings, correspondence, policies, fact summaries, exhibits), which you author yourself.
Run those tasks through frontier AI models and evaluate the deliverable they produce (the memo, redline, brief section or analysis) against the standard you would hold a colleague to.
Compare two model outputs on identical prompts and files, decide which performed better, and document where each fell short.
Write detailed grading rubrics that specify what a correct deliverable must contain, such as the right issues spotted, the right authorities applied, the right provisions flagged and the right register for the recipient, and explain in writing why a response passes or fails each one.
Flag concrete failures with evidence: fabricated or misread authorities, wrong standard applied, missed issues, provisions ignored, conclusions the facts do not support, and off brief interpretation of the ask.
Contribute across litigation, transactional, compliance and regulatory, and IP work, and review and refine tasks built by other experts.
2+ years legal experience preferred. A current practice is welcome but optional.
Depth in at least one of: litigation (civil, commercial, employment, or administrative); corporate and transactional (M&A, commercial contracts, financing); compliance and regulatory; intellectual property (patent, trademark, copyright); employment law; real estate; tax law; legal operations with sustained drafting responsibility.
Working understanding of several of the others, enough to know what those workflows involve and how they are run, so you can assess work in an adjacent practice area and point out what was done correctly or incorrectly.
Grounds conclusions in primary sources: statutes, regulations, case law, and the contract in front of you. Can tell a real authority from a plausible sounding fabricated one.
In progress Bachelor’s degree or higher. Law degree preferred (JD, LLB or equivalent). Bar admission is valued but not required; senior paralegals and legal analysts with sustained drafting experience are in scope. Bar admitted candidates may be asked to verify admission during the assessment.
US, UK, Canadian, Australian and New Zealand trained lawyers all convert well. Note the jurisdiction on referral.
2+ years of hands on experience in your field preferred (see Domain qualifications above). Candidates with less experience are considered where the practical work is real.
Able to draw on your own real world experience and day to day workflows to craft scenarios that test whether an AI system can actually do the work.
Hands on practitioner: you currently do (or recently did) the work yourself at an individual contributor level, not solely in a managerial capacity.
Full professional or native level written and spoken English, with strong written communication. You can explain complex professional reasoning clearly and concisely, and articulate why a result is wrong, not only that it is.
Comfort with ambiguity and attention to detail. You can orient in a new set of files and build an accurate, deep working picture of it quickly, especially when the subject sits partly outside your own specialization. You verify what a document claims against the underlying numbers, sources or facts.
Capable of interpreting feedback, judging which parts of it are actually correct, and applying it without hand holding. When stuck, you look for the answer rather than waiting for one.
Ability to ramp quickly on unfamiliar work from written material and instructions alone, including where that material is incomplete (for example, writing grading rubrics for the first time).
General familiarity with AI and LLM tools. You have used models like Claude or ChatGPT in professional work and have the judgment to tell a well reasoned answer from a plausible sounding but incorrect one.
Baseline tech literacy: comfortable with cloud file tools (e.g., Google Workspace), managing browser profiles, downloading and installing desktop apps (e.g., Claude), and everyday file handling (e.g., converting between Excel and Google Sheets, zipping files for sharing).
Available at least 10 hours per week, with no weekly maximum. Consistent availability is valued and full time hours are available.
Based in the United States, Canada, or the UK.