AI Evaluation Infrastructure Consultant

Landing Point

Village of Pelham (NY)

On-site

USD 152,000 - 179,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Landing Point seeks an AI Evaluation Infrastructure Consultant to ensure the quality and compliance of AI systems across production environments. The role builds evaluation disciplines and tooling, develops golden test sets, and integrates drift detection to support scalable governance.

In collaboration with risk, compliance, and legal teams, you will establish production-readiness gates, run regression and adversarial testing, and provide evaluation evidence for governance and third-party

Qualifications

  • 7+ years of experience in software quality, data science, machine learning, or related fields.
  • Experience designing evaluation methods for LLM or ML systems.
  • Strong understanding of testing methodology and statistical rigor.
  • Experience with bias and fairness evaluation; familiarity with fair-lending concepts is a plus.
  • Ability to define and enforce production-readiness gates and acceptance criteria.
  • Experience partnering with risk, compliance, and legal teams.

Responsibilities

  • Develop evaluation harnesses and tooling for AI systems.
  • Create golden test sets and scenario libraries for expected behaviors and edge cases.
  • Conduct regression testing to identify quality changes.
  • Perform hallucination and grounding/faithfulness testing.
  • Conduct bias and fairness testing to support fair-lending obligations.
  • Implement adversarial and red-team testing.
  • Ensure policy-adherence testing against compliance and regulatory requirements.
  • Monitor drift detection and ongoing production.
  • Manage human and subject-matter-expert evaluation workflows.
  • Establish production-readiness gates for AI systems.
  • Provide evaluation evidence and reporting for AI governance.
  • Define AI vendor acceptance criteria for third-party solutions.

Skills

LLM evaluation methods
Statistical testing
Bias and fairness evaluation
Production-readiness gates
Risk & compliance collaboration

Job description

Company Overview:

A leading financial services company is seeking an AI EvaluationInfrastructure Consultant toensure the quality and compliance of AI systems. This role is crucial for building the evaluation discipline and tooling necessary for AI systems to be accurate, safe, and ready to scale.

Job Responsibilities:
  • Develop evaluation harnesses and tooling for AI systems.
  • Create golden test sets and scenario libraries for expected behaviors and edge cases.
  • Conduct regression testing to identify quality changes.
  • Perform hallucination and grounding/faithfulness testing.
  • Conduct bias and fairness testing to support fair-lending obligations.
  • Implement adversarial and red-team testing.
  • Ensure policy-adherence testing against compliance and regulatory requirements.
  • Monitor drift detection and ongoing production.
  • Manage human and subject-matter-expert evaluation workflows.
  • Establish production-readiness gates for AI systems.
  • Provide evaluation evidence and reporting for AI governance.
  • Define AI vendor acceptance criteria for third-party solutions.
Qualifications:
  • 7+ years of experience in software quality, data science, machine learning, or related fields.
  • Experience designing evaluation methods for LLM or ML systems.
  • Strong understanding of testing methodology and statistical rigor.
  • Experience with bias and fairness evaluation; familiarity with fair-lending concepts is a plus.
  • Ability to define and enforce production-readiness gates and acceptance criteria.
  • Experience partnering with risk, compliance, and legal teams.
Compensation:

Pay Rate: $120/hr, DOE

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
AI Evaluation Infrastructure & Production Readiness Lead
AI Evaluation Infrastructure & Production Readiness Lead

Ports North • Baltimore (MD)

On-site
USD 138,000 - 207,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
AI Evaluation Expert - Remote
AI Evaluation Expert - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 76,000
AI Evaluation Specialist - Remote
AI Evaluation Specialist - Remote

YO AI Labs • New York (NY)

Remote
USD 34,000 - 55,000
AI Model Evaluation Lead: Metrics, Bias & Fairness
AI Model Evaluation Lead: Metrics, Bias & Fairness

MERIT Beauty • New York (NY)

On-site
Confidential
AI Evaluation & Reliability Architect
AI Evaluation & Reliability Architect

Landing Point • Village of Pelham (NY)

On-site
USD 152,000 - 179,000
Automation QA w/ AI testing
Automation QA w/ AI testing

Compunnel, Inc. • New York (NY), Northern (KY)

On-site
USD 140,000 - 190,000
AI QA Engineer
AI QA Engineer

Cavendish Professionals • Town of Italy (NY)

On-site
USD 95,000 - 120,000
AI Software Engineer – LLM Evaluation & Automation (Remote)
AI Software Engineer – LLM Evaluation & Automation (Remote)

Stage 4 Solutions Inc • United States

Remote
USD 99,000 - 108,000
Health benefits
401K