Senior Scientist - GenAI Evaluation

Johnson & Johnson Innovative Medicine

Madrid

Híbrido

EUR 55.000 - 88.000

Jornada completa

hace 26 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca la empresa.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Annual bonus
Vacation days
Parental leave

Descripción de la vacante

Johnson & Johnson is seeking a Senior Scientist in Generative AI Evaluation & Standards to design, build, and run evaluation methods across R&D. You will author rubrics, curate datasets, validate AI judges, and produce quality readouts guiding release decisions.

Hybrid role in Madrid or Barcelona with limited travel, 6+ years in AI/ML evaluation, strong Python skills, and experience with LLMs and RAG are preferred.

Formación

  • Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Biomedical Engineering, or a related field; PhD preferred.
  • 6+ years of hands-on experience in AI/ML evaluation or data science (Master's) or 3+ years with PhD.

Responsabilidades

  • Design, build, and maintain automated evaluation pipelines for AI/ML systems across R&D.
  • Author evaluation rubrics and scoring criteria, curate datasets, and maintain the registry of reusable evaluation assets.
  • Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks.
  • Analyze failure patterns and translate findings into actionable recommendations.
  • Develop evaluation criteria with scientific, clinical, and regulatory partners.
  • Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality.
  • Build evaluation tooling and reusable patterns for self-serve by other teams.

Conocimientos

Python
AI/ML evaluation
Data analysis
Collaboration

Educación

Master's degree
PhD preferred

Herramientas

LLM APIs
Vector databases
Embedding models
Evaluation harnesses

Descripción del empleo

At Johnson & Johnson,we believe health is everything. Our strength in healthcare innovation empowers us to build aworld where complex diseases are prevented, treated, and cured,where treatments are smarter and less invasive, andsolutions are personal.Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity.Learn more at jnj.com.
As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.

Job Function

Data Analytics & Computational Sciences

Job Sub Function

Data Science

Job Category

Scientific/Technology

All Job Posting Locations

Cornellà de Llobregat, Barcelona, Spain, Madrid, Spain

Job Description

At J&J we are building Generative AI solutions to support pharmaceutical R&D — literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support. These systems need to be evaluated before teams rely on them in scientific workflows. In pharma, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error — so quality must be measurable, repeatable, traceable, and scientifically defensible.

As a Senior Scientist in our Generative AI Evaluation & Standards (EQS) function within Data Science and Digital Health, you will design, build, and run the evaluation methods we depend on to assess these systems across R&D. You will author rubrics, curate benchmark and golden datasets, validate AI judges, and produce the quality readouts that inform release decisions. Want to shape how a top pharmaceutical company determines whether its AI is ready for real scientific work? This is that role!

Key Responsibilities

We need someone who can build the evaluation assets our teams count on, and continuously improve them based on what we learn. Day to day, you will:

  • Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy.
  • Author evaluation rubrics and scoring criteria, curate golden and synthetic datasets with domain experts, and maintain our registry of reusable evaluation assets.
  • Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks that produce standardized quality readouts.
  • Analyze failure patterns — hallucination, unsupported claims, weak traceability — and turn findings into actionable recommendations.
  • Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners, refining them based on real-world feedback.
  • Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality — where generic benchmarks fall short.
  • Build evaluation tooling and reusable patterns that enable other teams to self-serve.
Qualifications
Education

Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or a related field required. PhD preferred.

Required
EXPERIENCE AND SKILLS

We are looking for someone with 6+ years of hands-on experience in AI/ML evaluation or data science (Master's) or 3+ years of industry experience (PhD). You should have experience designing and running evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems, and hands-on work with generative AI — large language models, retrieval-augmented generation, agentic frameworks, and prompt engineering. We also value strong proficiency in Python and modern AI/ML tooling (evaluation harnesses, embedding models, vector databases, LLM APIs), the ability to translate expert scientific judgment into measurable criteria, rubrics, and reproducible protocols, and a collaborative, self-driven approach to working across multidisciplinary teams.

Preferred

We would love to find someone who also brings experience with AI/ML evaluation in regulated environments (FDA, EMA, or equivalent), understanding of the drug development pipeline and biomedical data types, or domain expertise in oncology, immunology, or neuroscience. Experience designing or validating LLM-as-judge systems, implementing CI/CD evaluation pipelines, or publications in AI evaluation, NLP, or biomedical informatics are all a plus.

Other

English proficiency is required (written and verbal). This is a hybrid role based in Madrid or Barcelona, with limited travel (<10%), primarily within Europe.

Required Skills
Preferred Skills

Advanced Analytics, Business Intelligence (BI), Coaching, Collaboration, Critical Thinking, Data Analysis, Database Management, Data Privacy Standards, Data Reporting, Data Savvy, Data Science, Data Visualization, Econometric Models, Process Improvements, Technical Credibility, Technologically Savvy, Workflow Analysis

The Anticipated Base Pay Range For This Position Is

€55,400.00 - €87,860.00

Benefits

In addition to base pay, we offer the following benefits*: an annual bonus with set target (% of pay) depending on pay grade / location, where the actual amount is based on the employees’ and companies’ performance of the previous calendar year, or sales commissions. Moreover, we offer vacation days, parental leave for a minimum of 12 weeks, bereavement leave, caregiver leave, volunteer leave, well-being reimbursement, programs for financial, physical and mental health. We also offer service anniversary and recognition awards, and subject to the terms of their respective plans, employees - and in some location’s eligible dependents - can participate in several insurance plans. For more information, visit Employee benefits | Supporting well-being & career growth | Johnson & Johnson Careers.

  • This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

Johnson & Johnson Co. • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+2
Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

Johnson & Johnson Innovative Medicine • Cornellà de Llobregat

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+3
Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

7300-Janssen-Cilag S.A. Legal Entity • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Associate Director, Evaluation and Quality Standards
Associate Director, Evaluation and Quality Standards

7300-Janssen-Cilag S.A. Legal Entity • Barcelona

Presencial
EUR 75.000 - 129.000
Annual bonus
Vacation days
Parental leave
+1
Senior Scientist - GenAI Evaluation
Senior Scientist - GenAI Evaluation

Johnson Johnson • Madrid

Híbrido
EUR 85.000 - 110.000
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins

jj • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+5
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins

Johnson & Johnson Co. • Madrid

Presencial
EUR 55.000 - 88.000
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins

Johnson & Johnson Innovative Medicine • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+2
Senior ML Engineer, Biologics Discovery - ES
Senior ML Engineer, Biologics Discovery - ES

jj • Madrid

Presencial
EUR 55.000 - 88.000
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins
Senior Scientist – Generative AI for Workflow, Applications and Plug-ins

7300-Janssen-Cilag S.A. Legal Entity • Madrid

Híbrido
EUR 55.000 - 88.000
Annual bonus
Vacation days
Parental leave
+2