Senior Scientist - GenAI Evaluation

Johnson Johnson

Madrid

Híbrido

EUR 85.000 - 110.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

Johnson & Johnson is seeking a Senior Scientist to lead the evaluation of Generative AI solutions for pharma R&D within the EQS function in Data Science and Digital Health. You will design, build, and run evaluation methods to ensure AI systems deliver quality, defensible results across literature review, evidence synthesis, and decision support.

You will author rubrics, curate benchmark datasets, validate AI judges, and produce comparable readouts to inform release decisions.

Formación

  • Master's degree in a related field required; PhD preferred.
  • 6+ years in AI/ML evaluation or data science (Master's) or 3+ years with a PhD.
  • Experience designing and running evaluation frameworks and benchmarks for AI/ML systems.
  • Hands-on work with generative AI, LLMs, retrieval-augmented generation, and prompt engineering.

Responsabilidades

  • Design, build, and maintain automated evaluation pipelines for LLM quality and RAG performance.
  • Author rubrics and curate datasets; maintain reusable evaluation assets.
  • Validate AI judges and run model, prompt, retriever, and agent benchmarks.
  • Analyze failure patterns and provide actionable recommendations.
  • Develop evaluation criteria with scientific and regulatory partners.
  • Design evaluation methods for scientific reasoning and evidence synthesis.
  • Build tooling to enable self-serve evaluation by other teams.

Conocimientos

AI/ML evaluation
Python
LLM APIs
Evaluation frameworks
Data science
Prompt engineering
Vector databases

Educación

Master's degree
PhD preferred

Herramientas

Evaluation harnesses
Embedding models
Vector databases
LLM APIs

Descripción del empleo

At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com .

As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.

Job Function:

Data Analytics & Computational Sciences

Job Sub Function:

Data Science

Job Category:

Scientific/Technology

All Job Posting Locations:

Cornellà de Llobregat, Barcelona, Spain, Madrid, Spain

Job Description:

At J&J we are building Generative AI solutions to support pharmaceutical R&D - literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support. These systems need to be evaluated before teams rely on them in scientific workflows. In pharma, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error - so quality must be measurable, repeatable, traceable, and scientifically defensible.

As a Senior Scientist in our Generative AI Evaluation & Standards (EQS) function within Data Science and Digital Health, you will design, build, and run the evaluation methods we depend on to assess these systems across R&D. You will author rubrics, curate benchmark and golden datasets, validate AI judges, and produce the quality readouts that inform release decisions. Want to shape how a top pharmaceutical company determines whether its AI is ready for real scientific work? This is that role!

KEY RESPONSIBILITIES:
  • Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy.
  • Author evaluation rubrics and scoring criteria, curate golden and synthetic datasets with domain experts, and maintain our registry of reusable evaluation assets.
  • Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks that produce standardized quality readouts.
  • Analyze failure patterns - hallucination, unsupported claims, weak traceability - and turn findings into actionable recommendations.
  • Develop therapeutic‑area‑specific evaluation criteria with scientific, clinical, and regulatory partners, refining them based on real‑world feedback.
  • Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality - where generic benchmarks fall short.
  • Build evaluation tooling and reusable patterns that enable other teams to self‑serve.
QUALIFICATIONS
Education:

Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or a related field required. PhD preferred.

EXPERIENCE AND SKILLS:
Required:

We are looking for someone with 6+ years of hands‑on experience in AI/ML evaluation or data science (Master's) or 3+ years of industry experience (PhD). You should have experience designing and running evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems, and hands‑on work with generative AI - large language models, retrieval‑augmented generation, agentic frameworks, and prompt engineering. We also value strong proficiency in Python and modern AI/ML tooling (evaluation harnesses, embedding models, vector databases, LLM APIs), the ability to translate expert scientific judgment into measurable criteria, rubrics, and reproducible protocols, and a collaborative, self‑driven approach to working across multidisciplinary teams.

Preferred:

We would love to find someone who also brings experience with AI/ML evaluation in regulated environments (FDA, EMA, or equivalent), understanding of the drug development pipeline and biomedical data types, or domain expertise in oncology, immunology, or neuroscience. Experience designing or validating LLM‑as‑judge systems, implementing CI/CD evaluation pipelines, or publications in AI evaluation, NLP, or biomedical informatics are all a plus.

OTHER:

English proficiency is required (written and verbal). This is a hybrid role based in Madrid or Barcelona, with limited travel (

Required Skills:
  • Advanced Analytics, Business Intelligence (BI), Coaching, Collaboration, Critical Thinking, Data Analysis, Database Management, Data Privacy Stan
Preferred Skills:
  • Advanced Analytics, Business Intelligence (BI), Coaching, Collaboration, Critical Thinking, Data Analysis, Database Management, Data Privacy Stan
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

Johnson & Johnson Innovative Medicine • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

7300-Janssen-Cilag S.A. Legal Entity • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

Johnson & Johnson Co. • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+2
Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

Johnson & Johnson Innovative Medicine • Cornellà de Llobregat

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+3
Associate Director, Evaluation and Quality Standards
Associate Director, Evaluation and Quality Standards

7300-Janssen-Cilag S.A. Legal Entity • Barcelona

Presencial
EUR 75.000 - 129.000
Annual bonus
Vacation days
Parental leave
+1
Sr Engineer – GenAI Quality Assurance
Sr Engineer – GenAI Quality Assurance

Tamarind Intelligence • Madrid

Híbrido
EUR 90.000 - 130.000
Sr Engineer – GenAI Quality Assurance
Sr Engineer – GenAI Quality Assurance

Johnson & Johnson Innovative Medicine • Madrid

Híbrido
EUR 70.000 - 110.000
Sr Engineer – GenAI Quality Assurance
Sr Engineer – GenAI Quality Assurance

Johnson & Johnson Innovative Medicine • Cornellà de Llobregat

Híbrido
EUR 70.000 - 110.000
Senior GenAI Evaluation Scientist — Pharma R&D Impact
Senior GenAI Evaluation Scientist — Pharma R&D Impact

Johnson & Johnson Innovative Medicine • Cornellà de Llobregat

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+3
Senior GenAI Evaluation Scientist: AI Quality & Standards
Senior GenAI Evaluation Scientist: AI Quality & Standards

Johnson & Johnson Innovative Medicine • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave