Data Scientist II

RXinsider LTD.

Philadelphia (Philadelphia County)

On-site

USD 72,000 - 119,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Elsevier is seeking a Data Scientist II to design and deploy AI systems that enhance clinical knowledge discovery and evidence extraction across vast scientific content. You will work with NLP, generative AI, and large-scale data to deliver reliable, production-ready insights for clinicians and researchers.

You will collaborate with cross-functional teams to turn complex clinical challenges into measurable outcomes, write robust Python code, and contribute scalable data pipelines for

Qualifications

  • Experience in data science, machine learning, artificial intelligence, NLP, statistics, applied mathematics, computer science, or a related quantitative area.
  • Experience with frontier LLMs such as OpenAI's GPTs, Anthropic's Claude, and Google's Gemini, including fine-tuning LLMs and/or SLMs.
  • Strong Python skills and a habit of writing clean, maintainable, well-tested code.
  • Solid grasp of ML fundamentals: supervised/unsupervised, feature engineering, model evaluation, selection, and performance measures.
  • Experience with structured, semi-structured, or unstructured data, especially large-scale text or clinical/scientific content datasets.
  • Familiarity with Pandas, NumPy, SciPy, Scikit-learn, PyTorch, TensorFlow, or Matplotlib.
  • Ability to translate complex requirements into practical, data-driven solutions with strong analytical skills.
  • Clear communication and collaboration with engineering, product, clinical, and business stakeholders.

Responsibilities

  • Design and build machine learning, NLP, and generative AI systems for clinical knowledge discovery, evidence extraction, decision support, and content understanding.
  • Work with large-scale, complex data including clinical documentation, scientific publications, datasets, knowledge graphs, ontologies, timestamps, and content across disciplines.
  • Apply techniques like classification, regression, clustering, ranking, feature engineering, deep learning, embeddings, LLMs, retrieval, and generative AI.
  • Develop capabilities for semantic search, information retrieval, entity extraction, content classification, recommendation, ranking, summarization, QA, and evidence-grounded generation.
  • Build, evaluate, fine-tune, prompt, and integrate models into production systems with quality, relevance, reliability, and clinical value.
  • Write clean, production-quality Python and contribute reusable DS components, pipelines for preprocessing, inference, experimentation, monitoring, and improvement.
  • Support deployment, monitoring, model maintenance, drift detection, and automated retraining of DS systems for clinicians and researchers.

Skills

Python
ML/NLP
Deep learning
Communication

Tools

Pandas
NumPy
SciPy
PyTorch
TensorFlow

Job description

What if your next AI model could help accelerate a medical breakthrough, uncover a critical scientific insight, or help researchers solve some of humanity's greatest challenges? At Elsevier, data science is about far more than algorithms and model performance. It is about applying advanced AI to help researchers, clinicians, educators, and institutions discover knowledge, assess evidence, generate insights, and advance science for the benefit of society.


Every day, millions of researchers rely on our products to navigate an ever-growing universe of scientific information. As a Data Scientist II, you will help build the intelligent systems that make scientific knowledge more discoverable, trustworthy, connected, and actionable.


This is AI with purpose. This is technology in service of scientific progress.


About The Role

In this role, you will design and build machine learning, NLP, and generative AI solutions that support clinical knowledge discovery, evidence extraction, decision support, and intelligent content understanding. You will work with large-scale clinical and scientific content and data, applying the right techniques to solve complex problems and deliver reliable, production-ready systems. Working closely with cross-functional partners, you will help turn ambiguous clinical and scientific challenges into measurable outcomes that improve how clinicians and researchers discover and apply knowledge.


About The Team

As part of a growing team of Data Scientists, you will take on some of the hardest problems in science. This team is building intelligent systems that can reason across scientific publications, research data, knowledge graphs, ontologies, metadata, taxonomies, citations, and content spanning every scientific discipline


Responsibilities


  • Design and build machine learning, NLP, and generative AI systems for clinical knowledge discovery, evidence extraction, decision support, and intelligent content understanding.

  • Work with large-scale, complex, and heterogeneous data, including clinical documentation, scientific publications, research datasets, knowledge graphs, ontologies, taxonomies, citations, metadata, and content spanning every clinical and scientific discipline.

  • Apply the right technique to each problem, using approaches such as classification, regression, clustering, ranking, feature engineering, deep learning, embeddings, LLMs, retrieval, and generative AI.

  • Develop capabilities for semantic search, information retrieval, entity extraction, content classification, recommendation, ranking, summarization, question answering, and evidence-grounded generation for clinical and scientific use cases.

  • Build, evaluate, fine-tune, prompt, and integrate models into robust production systems, while continuously improving quality, relevance, reliability, and clinical/user value.

  • Write clean, tested, production-quality Python and contribute reusable data science components, packages, and scalable data pipelines for preprocessing, inference, experimentation, monitoring, and continuous improvement.

  • Support deployment, monitoring, model maintenance, drift detection, automated retraining, and ongoing optimization of data science systems that clinicians and researchers depend on.

  • Collaborate with engineering, product, UX, analytics, research, clinical, and domain experts, and communicate technical concepts, model behavior, insights, trade-offs, and recommendations clearly to technical and non-technical audiences.


Requirements


  • Experience in data science, machine learning, artificial intelligence, NLP, statistics, applied mathematics, computer science, or a related quantitative area.

  • Experience working with frontier LLMs such as OpenAI's GPTs, Anthropic's Claude, and Google's Gemini, including fine-tuning LLMs and/or SLMs.

  • Strong Python skills and a habit of writing clean, maintainable, well-tested code.

  • A solid grasp of machine learning fundamentals, including supervised and unsupervised learning, feature engineering, model evaluation, model selection, and performance measurement.

  • Experience working with structured, semi-structured, or unstructured data, especially large-scale text or clinical/scientific content datasets.

  • Familiarity with common data science and machine learning tools such as Pandas, NumPy, SciPy, Scikit-learn, PyTorch, TensorFlow, or Matplotlib.

  • The ability to translate complex and ambiguous requirements into practical, measurable, data-driven solutions, with strong analytical thinking, problem-solving skills, and attention to quality.

  • Clear communication skills, a collaborative approach to working with engineering, product, clinical, and business stakeholders, and a genuine interest in building production-ready systems that improve health outcomes.


Why This Work Matters

Your models won't just process data, they'll help shape how clinicians document care, how students learn to practice medicine, and how researchers uncover insights that improve patient outcomes. This is AI applied where it counts.


Work in a Way That Works for You

We promote a healthy work/life balance across the organisation. We offer an appealing working prospect for our people. With numerous wellbeing initiatives, parental leave, and study assistance, we will help you meet your immediate responsibilities and your long-term goals.


U.S. National Base Pay Range: $71,600 - $119,400. Geographic differentials may apply in some locations to better reflect local market rates. This job is eligible for an annual incentive bonus.


We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist II
Data Scientist II

Elsevier • Philadelphia

On-site
USD 72,000 - 119,000
Senior Data Scientist
Senior Data Scientist

RELX • Philadelphia

On-site
USD 95,000 - 159,000
Senior Data Scientist
Senior Data Scientist

RXinsider LTD. • Philadelphia

On-site
USD 95,000 - 191,000
Senior Data Scientist
Senior Data Scientist

Elsevier • Philadelphia

Hybrid
USD 95,300 - 158,800
Annual incentive bonus
Work with a rich collection of scientific data
Manager Data Science
Manager Data Science

RXinsider LTD. • Philadelphia

On-site
USD 115,000 - 192,000
Manager Data Science
Manager Data Science

RELX • Philadelphia

On-site
USD 115,000 - 231,000
Annual incentive bonus
Country specific benefits
Manager Data Science
Manager Data Science

LexisNexis Risk Solutions • Philadelphia

On-site
USD 115,000 - 193,000
Annual incentive bonus
Sr Data Scientist
Sr Data Scientist

RXinsider LTD. • Philadelphia

On-site
USD 120,000 - 180,000
Manager Data Science
Manager Data Science

Elsevier • Philadelphia

On-site
USD 115,400 - 192,300
Annual incentive bonus
Country-specific benefits
Senior AI Data Scientist I
Senior AI Data Scientist I

Exelixis • Alameda (CA)

On-site
USD 143,000 - 203,000
401(k) with company contributions
Health, dental, vision
Life and disability insurance
+2