Artificial Intelligence Researcher

Verita AI

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Verita AI is seeking an Applied AI Researcher to work with clients on model evaluation and data strategy. You will assess model performance, identify failure modes, and design data-driven solutions, collaborating with operations and engineering to implement scalable data programs.

The role requires strong Python skills, experience with model APIs, and the ability to communicate findings to clients and stakeholders clearly.

Qualifications

  • Experience in applied AI research, machine learning, model evaluation, or data-centric AI.
  • Experience evaluating foundation models or generative AI systems.
  • Strong Python skills and experience working with model APIs and structured datasets.
  • Ability to translate model failures into practical data solutions.
  • Strong technical writing and client communication skills.

Responsibilities

  • Work with clients to understand their models, goals, and performance gaps.
  • Design evaluations for generative, multimodal, reasoning, tool-use, and agentic AI systems.
  • Analyze model outputs and benchmark results to identify and quantify failure modes.
  • Recommend data solutions such as supervised fine-tuning data, preference data, expert demonstrations, critiques, and evaluation datasets.
  • Write client proposals covering the methodology, data design, quality controls, staffing, deliverables, and expected impact.
  • Create annotation guidelines, scoring rubrics, gold-standard tasks, and evaluator-training programs.
  • Design pilot studies and measure whether data interventions improve model performance.
  • Build quality systems using calibration tasks, blind review, adjudication, and expert scoring.
  • Work with operations and engineering teams to launch and scale data pipelines.
  • Present findings and recommendations to clients.

Job description

Verita AI works with leading AI companies to identify model gaps and build the human data needed to improve model performance. Verita AI operates a vetted expert network that connects specialized professionals with leading AI laboratories and human-data companies. The network comprises more 5,000 experts across finance, medicine, law, engineering, music, and other professional domains.

We recently raised a $6 million seed round led by Kindred Ventures.

About the Role

We are hiring an Applied AI Researcher to work directly with clients on model evaluation and data strategy.

You will evaluate model performance, identify failure modes, and recommend the datasets, rubrics, expert workflows, and quality controls needed to address them. You will then work with our operations and engineering teams to turn these recommendations into scalable data programs.

What You’ll Do
  • Work with clients to understand their models, goals, and performance gaps.
  • Design evaluations for generative, multimodal, reasoning, tool-use, and agentic AI systems.
  • Analyze model outputs and benchmark results to identify and quantify failure modes.
  • Recommend data solutions such as supervised fine-tuning data, preference data, expert demonstrations, critiques, and evaluation datasets.
  • Write client proposals covering the methodology, data design, quality controls, staffing, deliverables, and expected impact.
  • Create annotation guidelines, scoring rubrics, gold-standard tasks, and evaluator-training programs.
  • Design pilot studies and measure whether data interventions improve model performance.
  • Build quality systems using calibration tasks, blind review, adjudication, and expert scoring.
  • Work with operations and engineering teams to launch and scale data pipelines.
  • Present findings and recommendations to clients.
What We’re Looking For
  • Experience in applied AI research, machine learning, model evaluation, or data-centric AI.
  • Experience evaluating foundation models or generative AI systems.
  • Strong understanding of benchmark design, human evaluation, rubric development, and statistical analysis.
  • Ability to translate model failures into practical data solutions.
  • Strong Python skills and experience working with model APIs and structured datasets.
  • Familiarity with supervised fine-tuning, preference optimization, RLHF/RLAIF, reward modeling, synthetic data, or LLM-as-a-judge evaluation.
  • Strong technical writing and client communication skills.
  • Ability to independently structure and execute ambiguous research projects.
Nice to Have
  • Experience at an AI lab, foundation-model company, AI data company, or post-training team.
  • Experience designing expert-data or human-evaluation programs.
  • Experience evaluating multimodal, coding, agentic, or tool-use systems.
  • Publications at conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP.
  • Previous client-facing research, consulting, solutions engineering, or forward-deployed experience.
  • Public research, code, benchmarks, or evaluation frameworks.

Please make sure you have any relevant work samples, including model evaluations, benchmarks, error analyses, research, technical writing, or code repositories.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Researcher: Model Evaluation & Data Strategy
Applied AI Researcher: Model Evaluation & Data Strategy

Verita AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Applied AI Researcher
Applied AI Researcher

Morpheus Talent Solutions • San Francisco (CA)

Hybrid
USD 200,000 - 350,000
Applied AI Researcher (Dublin, CA)
Applied AI Researcher (Dublin, CA)

Articul8 AI • Dublin (CA)

On-site
USD 120,000 - 160,000
Applied AI Researcher (Brazil)
Applied AI Researcher (Brazil)

Articul8 • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Research, Post-Training Evals
Research, Post-Training Evals

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Research, Post-Training Evals
Research, Post-Training Evals

Mosaic.tech • San Francisco (CA)

On-site
USD 140,000 - 190,000
AI Research Scientist
AI Research Scientist

Elios Talent • Austin (TX)

On-site
USD 120,000 - 150,000
Applied Research Engineer
Applied Research Engineer

Sterling Inspired Staffing. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive medical plans
Generous parental leave
Unlimited PTO
+3
Evaluation Lead
Evaluation Lead

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity