Director AI Evaluation

Geisinger

Pennsylvania

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Remote work

Job summary

Geisinger is seeking a Director of AI Evaluation to define and sustain quality across its AI portfolio, including internal models and vendor systems. You will lead a team of data scientists and senior analysts, setting standards and ensuring rigorous evaluation from design through production.

The role blends people leadership with technical authority, guiding major AI programs and reporting findings to the VP of AI and executive leaders to drive measurable clinical outcomes and value for

Qualifications

  • Bachelor’s Degree in a related field (Required).
  • 8+ years of related work experience (Required).
  • 3+ years of managerial/supervisory experience (Required).

Responsibilities

  • Lead AI evaluation programs and set quality standards across internal and vendor AI systems.
  • Manage data science and evaluation teams, providing hands-on guidance on validation and equity audits.
  • Translate failures into production monitoring metrics and collaborate with the AI Platform team.
  • Report findings to the VP of AI and senior leadership.

Skills

People leadership
Experimental design
Model evaluation in production
LLM / generative AI evaluation
Python
SQL
Fairness in ML
Communication
Healthcare domain knowledge

Education

Bachelor's degree in Related Field

Job description

Overview

Location: Work from home (Pennsylvania). Shift: Days (United States of America). Scheduled Weekly Hours: 40. Exemption Status: Yes. Job summary: The Director of AI Evaluation owns how Geisinger defines, proves, and sustains quality across its AI portfolio, including internally built models and vendor-provided systems. This hands-on technical leader also manages the people who build and validate AI, including data scientists and senior analysts who evaluate AI systems. The Director sets the standard for quality, leads the team that enforces it, and reports findings to the VP of AI, executive leaders, and the board. This role combines people leadership with technical authority to guide major AI programs toward evidence that withstands scrutiny.

Responsibilities
  • Reports to the VP of AI; directly manages the data science line and matrix-manages the Senior Analysts, AI Evaluation.
  • Determine the quality standard for high-value AI initiatives (internal or vendor-provided) from design through production; hold both built and bought systems to the same standard.
  • Own the methodology for pre-production validation and live production monitoring.
  • Lead the health and development of the data science and evaluation teams, providing hands-on technical guidance on validation studies, equity audits, monitoring plans, and escalation playbooks.
  • Maintain the evaluation toolkit, reusable playbooks, and templates to accelerate new programs.
  • Translate failure modes into concrete, measurable production-monitoring metrics; collaborate with the AI Platform team on backend implementation.
  • Track AI system performance against clinical thresholds; monitor user adoption, engagement, and time-to-action; distinguish genuine value from alarm fatigue.
  • Link each AI initiative to the outcomes it aims to improve (e.g., mortality, time-to-treatment, boarding time, denial rate, cost per case) using pre-launch baselines and appropriate horizons.
  • Monitor equity by assessing performance gaps across relevant subgroups to surface disparate impact early.
  • Ensure compliance with organizational policies and procedures; perform duties aligned with job obligations.
Qualifications
  • People-leadership experience — managing, developing, and growing technical staff; building teams.
  • Strong foundation in experimental design and causal inference; ability to select appropriate methods for given situations.
  • Hands-on experience designing and running model evaluation studies in production settings.
  • Experience evaluating LLM or generative AI systems, or comparable complex ML systems with ambiguous or noisy ground truth.
  • Ability to translate ambiguous failure modes into concrete, defensible evaluation designs and monitoring metrics.
  • Strong fluency in Python and SQL; comfortable with modern ML tooling and cloud-native data environments.
  • Experience evaluating fairness and equity in ML systems.
  • Clear written communication for evaluation memos and specifications used by non-technical decision-makers.
  • Healthcare, clinical, or regulated-industry experience strongly preferred.
Education & Experience
  • Bachelor’s Degree in Related Field (Required).
  • Minimum 8 years related work experience (Required).
  • Minimum 3 years managerial/supervisory experience (Required).
Values
  • OUR PURPOSE & VALUES: Everything we do is about caring for our patients, our members, our students, our Geisinger family and our communities.
  • KINDNESS: Treat everyone as we would hope to be treated ourselves.
  • EXCELLENCE: Strive for excellence in all we do.
  • LEARNING: Share knowledge to prepare caregivers for tomorrow.
  • INNOVATION: Seek new and better ways to care for patients and the community.
  • SAFETY: Provide a safe environment for patients and members.

We offer healthcare benefits for full-time and part-time positions from day one, including vision and dental coverage. We value collaboration, cooperation, and collegiality and believe a diverse workforce strengthens our team. We are an affirmative action, equal opportunity employer. All qualified applicants will receive consideration for employment regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status. For more information, visit www.geisinger.org or connect with us on social media.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead Data Scientist, AI Evaluation & Monitoring
Tech Lead Data Scientist, AI Evaluation & Monitoring

Geisinger • Danville (PA)

Hybrid
USD 120,000 - 150,000
Senior BI Analyst: AI Evaluation & Monitoring (Remote)
Senior BI Analyst: AI Evaluation & Monitoring (Remote)

Geisinger • Danville (PA)

Hybrid
USD 85,000 - 110,000
Data Scientist Team Lead
Data Scientist Team Lead

Socket.dev • Pennsylvania

Hybrid
USD 120,000 - 170,000
Senior BI Analyst, AI Discovery & Strategy
Senior BI Analyst, AI Discovery & Strategy

Geisinger • Danville (PA)

On-site
USD 85,000 - 110,000
Remote Director of AI Quality & Evaluation
Remote Director of AI Quality & Evaluation

Geisinger • Pennsylvania

Hybrid
USD 180,000 - 240,000
Health benefits
Remote work
AI Data Scientist Senior
AI Data Scientist Senior

Geisinger • Danville (PA)

On-site
USD 90,000 - 130,000
Healthcare benefits from day one
Opportunities for professional growth
Diversity and inclusion initiatives
Data Scientist Team Lead
Data Scientist Team Lead

3M HEALTHCARE • Town of Montana (WI)

Hybrid
USD 140,000 - 190,000
Healthcare benefits from day one
Vision benefits
Dental benefits
+1
Business Intelligence Analyst Senior (AI Evaluation & Monitoring)
Business Intelligence Analyst Senior (AI Evaluation & Monitoring)

Geisinger • Danville (PA)

On-site
USD 85,000 - 110,000
Senior AI Product Manager
Senior AI Product Manager

Geisinger • United States

Hybrid
USD 120,000 - 160,000
Senior Digital Solutions Analyst – Analytics & AI
Senior Digital Solutions Analyst – Analytics & AI

Geisinger • Danville (PA)

On-site
USD 90,000 - 120,000
Healthcare benefits