Get more replies from employers
Send a job-specific resume in minutes.
Mercor, in partnership with a leading AI research organization, seeks experienced accountants to define task-specific grading criteria for real-world accounting deliverables and to score both AI-generated and human samples with rigorous, written explanations.
The role emphasizes deep knowledge of US GAAP, month-end close, reconciliations, and financial statement preparation, with a strong emphasis on precision, consistency, and defensible judgments.
Mercor is partnering with a leading AI research organization to engage experienced accountants for a project focused on evaluating how well AI systems perform real-world accounting work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.
Design precise, task-specific grading criteria for real-world accounting deliverables (reconciliations, close packages, financial statements and disclosures, workpapers, technical accounting memos)
Score AI-generated and human work samples against those criteria, with detailed written justifications for every score
Apply consistent, evidence-based judgment so that scores are reproducible and defensible
Incorporate structured feedback from senior reviewers and iterate quickly on your work
5+ years of professional accounting experience
Background in audit, assurance, or advisory at leading global accounting firms, or in senior accounting roles at large public companies
CPA or equivalent professional certification strongly preferred
Deep fluency in the day-to-day craft: US GAAP application, month-end close, account reconciliations, financial statement preparation and disclosures, and audit-ready workpapers
Exceptionally strong written communication, with the ability to explain exactly why a piece of work does or does not meet professional standards
Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers
Prior experience with AI training, evaluation, or human-data projects is a strong plus