Lead, Legal AI Benchmarking & Evaluation

Newcode.ai

New York (NY)

On-site

USD 120,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Newcode.ai in New York is seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform, ensuring our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate.

You will design benchmarks for contract drafting, information extraction, legal research,

Qualifications

  • Strong statistical background and analytical mindset.
  • Experience designing evaluation benchmarks for AI/ML models.

Responsibilities

  • Design Legal Benchmarks for: Contract Drafting, Information Extraction, Legal Research, and Contract Review
  • Build, source and maintain relevant datasets
  • Audit AI Output: Review and score complex AI-generated legal text, contract analyses, and statutory interpretations for accuracy and precision and lay out a strategy.
  • Define Evaluation Metrics: Establish clear criteria for grading model performance, specifically focusing on logical reasoning, citation accuracy, and the model's ability to safely abstain from answering.
  • Collaborate with Engineering: Partner directly with Engineering to translate legal errors into actionable technical feedback for model fine-tuning

Skills

Analytical skills
Statistical background
Attention to detail

Education

PhD or Masters in statistics, mathematics, ML or equivalent

Tools

Python
Pandas
NumPy
Jupyter notebooks

Job description

Newcode.ai in New York is seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform, ensuring our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate.

You will design benchmarks for contract drafting, information extraction, legal research,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of Legal AI Benchmarking & Evaluation
Head of Legal AI Benchmarking & Evaluation

Newcode.ai • Town of Sweden (NY)

On-site
USD 90,000 - 130,000
Head of Legal AI Evaluation & Benchmarking
Head of Legal AI Evaluation & Benchmarking

Newcode.ai • United States

On-site
USD 120,000 - 180,000
Head of Evaluations (Legal AI Benchmarking)
Head of Evaluations (Legal AI Benchmarking)

Newcode.ai • New York (NY)

On-site
USD 120,000 - 190,000
Head of Evaluations (Legal AI Benchmarking)
Head of Evaluations (Legal AI Benchmarking)

Newcode.ai • Town of Sweden (NY)

On-site
USD 90,000 - 130,000
Head of Evaluations (Legal AI Benchmarking)
Head of Evaluations (Legal AI Benchmarking)

Newcode.ai • United States

On-site
USD 120,000 - 180,000
AI-Driven Legal Solutions Architect
AI-Driven Legal Solutions Architect

Newcode • New York (NY)

On-site
USD 80,000 - 120,000
Senior Legal AI Specialist — Benchmark & QA Lead
Senior Legal AI Specialist — Benchmark & QA Lead

Mercor • New York (NY)

Hybrid
USD 180,000 - 240,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
AI-Driven Legal Engineer
AI-Driven Legal Engineer

Nextlex • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 160,000