GSK, we have bold ambitions for patients, aiming to positively impact the health of 2.5 billion people by the end of the decade. Our R&D focuses on discovering and delivering vaccines and medicines, combining our understanding of the immune system with cutting-edge technology to transform people's lives. GSK fosters a culture ambitious for patients, accountable for impact, and committed to doing the right thing, making sure that we focus our efforts on accelerating significant assets that meet patients' needs and have the highest probability of success. We're uniting science, technology, and talent to get ahead of disease together.
Find out more: Our approach to R&D.
Position Summary:
In this role you will act as the technical engine of the Analytics & AI pillar. The primary purpose of the role is to leverage AI to automate R&D Master Data (Project, Product and Substance) related processes, and predictive analytics to derive insights from data. The role also requires analyzing data, drawing appropriate conclusions and deriving insights. This role acts as the bridge between our team and Tech, ensuring to anchor our requirements into reality, and feasibility.
Responsibilities:
Advanced Data Preparation & Analysis (Main Focus):
- Utilize Python/PySpark and SQL to extract, clean, and structure complex R&D master data from multiple enterprise systems (e.g., ERP, PLM, RIM) and perform necessary transformations
- Perform deep-dive data profiling to identify hidden anomalies, root causes of data friction, and recurring quality issues within Project, Product, and Substance domains.
AI Development & Analysis (Main Focus):
- Performing feature engineering, structuring training datasets, and writing Python scripts required for the development of AI PoCs.
- Gather business requirements, perform analysis of requirements, translate them into technical requirements.
- Assist in deploying lightweight predictive algorithms (e.g., duplicate detection scripts) directly into the Power BI ecosystem or through upstream data preparation layers
Development & Visualization (Support Focus):
- Assist in designing, building, and deploying interactive Power BI dashboards that serve as the daily operational radar for the E2E Data Stewardship Hub and R&D business leaders.
- Translate qualitative feedback from Data Stewards and Coordinators into tangible dashboard features that actively reduce their daily manual workload.
- Assist in deploying lightweight predictive algorithms (e.g., duplicate detection scripts) directly into the Power BI ecosystem or upstream data preparation layers.
Compliance & Data Governance:
- Ensure all analytical models, scripts, and BI dashboards adhere to enterprise data governance policies and GxP documentation standards.
Ability to support MDM Experts:
- Ensure ability and capacity to perform the tasks of the MDM experts to support when needed, but also to lead the automation of these tasks with greater precision and independence.
Why You?
Work arrangement: This role is based in Poland and is offered on a hybrid working basis. You will be expected to work on-site part of the week in line with team and business needs.
Basic Qualifications & Skills:
- At least bachelor's degree in Data Science, Computer Science, Statistics, Mathematics, Physics, Engineering, a Life Sciences discipline with a strong quantitative focus or related field
- 3 to 5 years of professional experience in data analytics, BI development, or data engineering, with proven experience in implementing AI solutions
- Experience in implementing projects & solutions using Microsoft Azure stack, must have are Databricks
- Python, PySpark and SQL programming skills
- Technical Translation: Ability to translate complex scientific data requirements into clear, executable technical specifications for IT developers.
- Knowledge of GitHub
- Analytical Problem Solving: A proactive mindset capable of looking at a messy dataset and independently figuring out how to structure it for both reporting and future machine learning.
- Agile Collaboration: Comfortable working in short, iterative sprints alongside AI engineers, Process Architects, and business stakeholders.
- Fluent English skills
Preferred Qualifications & Skills:
- Experience with implementing GenAI solutions in Databricks
- Experience with Langchain
- Experience working with master data or R&D data within the Pharmaceutical, Biotech, or Vaccines industry.
- Industry Knowledge: Strong understanding of pharmaceutical data standards, GxP regulations, and ideally exposure to ISO IDMP frameworks.
- Fluency in Polish language
Skills
Artificial Intelligence (AI), Computational Sciences, Computer Programming, Data Analysis, Data Management, Data Modeling, Data Science, Predictive Modeling, Statistical Models
Polish Salary Range / Polski przedział wynagrodzenia: PLN 228,000 to PLN 380,000The annual gross base salary range for new hires in this position is listed above for each applicable location. These ranges take into account a number of f