Tech Lead Data Scientist, AI Evaluation & Monitoring

Geisinger

Danville (PA)

Hybrid

USD 120,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Geisinger seeks a Tech Lead Data Scientist in Danville, PA, to lead the evaluation and optimization of AI systems in production. The role involves mentoring data analysts, designing validation studies, and using advanced methodologies to ensure high-quality AI monitoring. Candidates should have 6+ years in data science and strong fluency in Python and SQL, ideally with healthcare experience. This position offers the chance to work in a hands-on technical leadership role within a collaborative team.

Qualifications

  • 6+ years in data science, statistics, ML engineering, or applied quantitative research.
  • Strong foundation in experimental design and causal inference.
  • Experience evaluating LLM or generative AI systems.

Responsibilities

  • Lead evaluation methodology for AI programs across the enterprise.
  • Provide hands-on guidance for the design of validation studies.
  • Mentor and set technical direction for a team of data analysts.

Skills

Data science
ML engineering
Experimental design
Causal inference
Python
SQL
Fairness evaluation
Technical leadership

Education

Bachelor's Degree in a related field
MS or PhD in a quantitative field (preferred)

Tools

Python
SQL
Cloud-native data platforms

Job description

Job Summary

The Tech Lead Data Scientist, AI Evaluation & Monitoring is the principal technical expert for how Geisinger evaluates, monitors, and optimizes AI systems in production. This hands‑on technical leadership role sets the technical direction for AI evaluation across a large portfolio, provides leadership to a team of data analysts, and partners directly with AI program teams to raise the quality of AI validation, monitoring, and improvement.

Job Duties
  • The technical evaluation methodology applied to AI programs across the enterprise – pre‑production validation, production monitoring, and ongoing optimization.
  • Hands‑on guidance to program teams as they design validation studies, equity audits, monitoring plans, and escalation playbooks.
  • Instrumentation of production monitoring: translating program‑specific failure modes into concrete, measurable metrics.
  • The evaluation toolkit: LLM-as-Judge frameworks, golden sets, simulation harnesses, experimental study designs, drift detection, subgroup fairness analysis.
  • Reusable evaluation playbooks and templates that let each new program move faster.
  • Technical direction, design review, and mentorship for a team of data analysts supporting the evaluation function.
What You Will Not Own
  • People management, HR administration, or formal performance evaluations for the analyst team.
  • Program‑level product strategy or go/no‑go decisions.
  • Final clinical validation judgment on whether a given AI is safe for a specific clinical use.
  • The software infrastructure behind the evaluation and monitoring tooling.
Shape Of The Work

With program teams (hands‑on advisory). Partner with program owners early to shape study approach, sample size, stratification, gold‑standard definition, and decision thresholds. Translate ambiguous failure modes into concrete, defensible evaluation designs and coach teams through technical work.

With the evaluation toolkit (hands‑on build). Design and operate reusable assets that let evaluation scale: LLM‑as‑Judge rubrics and calibration methods, golden sets, simulation harnesses, A/B and shadow‑mode study templates, subgroup fairness analyses, and drift monitors.

With the analyst team (technical leadership). Set technical direction, assign work across active evaluations, review analysis code and study designs, and raise the technical bar. Mentor analysts on methodology, statistical rigor, and domain knowledge.

Methods You'll Use
  • Experimental and quasi‑experimental design for production AI systems.
  • LLM and generative AI evaluation: golden sets, judge‑based evaluation, hallucination and grounding checks.
  • Fairness and equity evaluation across patient and stakeholder subgroups.
  • Production monitoring design: drift detection, performance decay, adoption, and outcome metrics.
  • Causal inference methods appropriate to healthcare settings where full RCTs are impractical.
  • Simulation and adversarial testing for pre‑production stress testing.
  • Python, SQL, modern ML and evaluation tooling, cloud‑native data platforms.

Work is typically performed in an office or remote environment and requires compliance with all organization policies and procedures.

Required Skills & Qualifications
  • 6+ years in data science, statistics, ML engineering, or applied quantitative research, with senior technical voice on cross‑functional projects.
  • Strong foundation in experimental design and causal inference.
  • Hands‑on experience designing and running model evaluation studies in real production settings.
  • Experience evaluating LLM or generative AI systems, or comparable complex ML systems where ground truth is messy.
  • Proven ability to translate ambiguous failure modes into concrete, defensible evaluation designs and monitoring metrics.
  • Strong fluency in Python and SQL; comfort with modern ML tooling and cloud‑native data environments.
  • Experience with fairness and equity evaluation for ML systems.
  • Track record of providing technical leadership and mentorship without formal people‑management authority.
  • Clear written communication – produces evaluation memos and specifications relied upon by non‑technical decision‑makers.
  • Healthcare, clinical, or regulated‑industry experience strongly preferred.
  • MS or PhD in a quantitative field preferred; equivalent experience accepted.
Education

Bachelor's Degree – Related Field of Study (Required)

Experience

Minimum of 6 years – Relevant experience (Required)

We are proud to be an affirmative action, equal‑opportunity employer, and all qualified applicants will receive consideration for employment regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Business Intelligence Analyst Senior (AI Evaluation & Monitoring)
Business Intelligence Analyst Senior (AI Evaluation & Monitoring)

Geisinger • Danville (PA)

On-site
USD 85,000 - 110,000
AI Data Scientist Senior
AI Data Scientist Senior

Geisinger • Danville (PA)

On-site
USD 90,000 - 130,000
Healthcare benefits from day one
Opportunities for professional growth
Diversity and inclusion initiatives
Senior BI Analyst: AI Evaluation & Monitoring (Remote)
Senior BI Analyst: AI Evaluation & Monitoring (Remote)

Geisinger • Danville (PA)

Hybrid
USD 85,000 - 110,000
Data Scientist (AI Quality & Evaluation)
Data Scientist (AI Quality & Evaluation)

Bioscope.ai, Inc. • Boston (MA)

On-site
USD 100,000 - 130,000
Data Scientist Team Lead
Data Scientist Team Lead

3M HEALTHCARE • Town of Montana (WI)

On-site
USD 140,000 - 190,000
Healthcare benefits from day one
Vision benefits
Dental benefits
+1
AI Engineer
AI Engineer

Harnham • San Francisco (CA)

On-site
USD 100,000 - 150,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
Director of AI & Machine Learning
Director of AI & Machine Learning

Accentuate Staffing • Morrisville (NC)

On-site
USD 120,000 - 150,000
Senior AI Engineer (JR229574)
Senior AI Engineer (JR229574)

ViziRecruiter,LLC. • City of Yonkers (NY)

On-site
USD 150,000 - 230,000
AL/ML Technical Lead
AL/ML Technical Lead

Booz Allen Hamilton • Atlanta (GA)

On-site
USD 128,700 - 292,000