Valvian is a Portugal-based biotechnology company developing next-generation therapeutics across oncology and age-related disorders.
From Lisbon, we are building an integrated biotech platform that combines disease biology, translational biomarkers, advanced therapeutic modalities, and computational discovery.
Beyond Valvian’s known programmes in oncology and age-related disorders, the successful candidate will also contribute to a confidential research initiative in a field we are helping define. Specifics will be shared under NDA after an initial conversation.
About the Role
You would be the first scientist in Valvian's Computational Biology & AI function, working directly with the Head of Computational Biology & AI.
We are building a computational platform that brings together AI models and proprietary experimental data to find the targets and drug candidates in oncology and age-related disorders that the existing knowledge does not make obvious.
The first major challenge sits upstream of any AI model. In our therapeutic area, the evidence is scattered across different public resources, primary literature and internal experimental data, reported inconsistently, and deeply dependent on experimental context that is rarely recorded. Whether a data substrate is structured, provenance-aware and honest about uncertainty determines whether anything built on top of it can be trusted. A model trained on unstructured data across the literature learns the shape of that literature and its biases rather than the underlying biology.
That structured data substrate is the core of the role. What should the data model look like? Which observations can legitimately be compared? How should conflicting evidence be represented rather than quietly reconciled? How are different entities dependent on each other? How do we unify existing data with our own? You would help answer those and other questions, then implement the pipelines and processes to populate, maintain and expand this resource as it becomes a reliable substrate for the implementation of predictive AI models.
Because the function is new, the work is unusually broad. You would have active participation in developing the full platform, i.e. the data model, the pipelines, the curation, the automations, and the conversations with the bench, rather than owning one narrow slice of an existing system.
Key Responsibilities
Data foundations
- Take raw, messy experimental and public data through to something usable: ingestion, cleaning, structuring, and the judgement calls that go with it.
- Build and maintain the pipelines and tools the rest of the team depends on, with version control, environment management, reproducibility and testing where it matters from the start.
- Work in modern cloud infrastructure, with attention to GDPR compliance and data sovereignty from day one.
Evidence model and curation
- Help design the data model that holds our evidence: the entities, relationships, controlled vocabularies and confidence scales that let a claim be assessed rather than assumed, and then curate against it at volume.
- Integrate heterogeneous data while preserving experimental context, provenance and uncertainty, so that what a measurement actually means survives the act of storing it.
- Systematically evaluate public biological databases against a common specification, establishing what each can and cannot support, and what integrating it would genuinely cost.
- Surface contradictions between sources rather than resolving them silently.
Automated and agentic workflows
- Design and evaluate semi-automated extraction of evidence from scientific literature, including deciding how such a system should be scored and where it should not be trusted.
- Treat model outputs as scientific inputs rather than finished answers, with explicit measurement of where they succeed and fail.
- Help build the first predictive models on top of the evidence layer, and design validation that tests whether they generalise to cases they have not seen rather than reproducing what they were shown.
Scientific partnership
- Work directly with our lab scientists on what their data records and what it does not, in the language of the experiment rather than through an interpreter.
- Help translate computational findings into experiments worth running, using what the platform reveals as uncertainty to identify where a measurement would be most informative.
- Provide a second pair of eyes on domain claims before they inform decisions, including ours. Saying that a conclusion does not follow is part of the job.
- Produce documentation clear enough that the rest of the team can understand and build on what you have made.
What you need to have
- Three to six years in computational biology, bioinformatics, proteomics, structural biology or a related quantitative field. A PhD or strong hands‑on experience are preferred.
- Experience taking raw, messy biological data through a working analysis or tool, start to finish, generating meaningful, robust and traceable results.
- Strong Python: modules, tests, version control, code someone else can run six months later. SQL, R and workflow tools such as Nextflow, Snakemake or Airflow are all useful here.
- Something you built that other people and/or systems have used: a pipeline, a database, a process, a tool. Even better if you have experience building and maintaining data ingestion and harmonisation pipelines.
- Strong experience in LLMs and agentic systems in the context of scientific work.
- The judgement to tell an observation from an interpretation from a hypothesis, in other people’s work and in your own. Reading a methods section and saying what it establishes, what it does not and what is missing is central to the role.
- Comfort working from first principles with incomplete specification, and the instinct to ship a rough version to learn from rather than wait for a full picture.
- The ability to work across the computational and experimental boundary.
Any of these would strengthen your application:
- Mass spectrometry, proteomics or structural biology workflows; affinity reagents and their validation, including antibody specificity, binding assays and enzymatic controls; ontologies, knowledge graphs or structured data integration; uncertainty‑aware modelling, active learning or experimental design; literature mining, LLM‑based extraction of structured data from papers, or any agentic system you built and measured; a public track record of code we can look at; prior biotech, drug discovery or translational research.
- Deep expertise in our specific biology is not required.
- Discovery orientation: comfortable working from first principles with incomplete specification. Enjoys the phase where requirements emerge through building and learning, rather than executing pre‑defined deliveries.
- Critical thinking: challenges briefs, tooling, and methods with evidence. Does not default to consensus.
- Ownership: drives decisions end‑to‑end, from definition to validation. Owns outcomes, not just deliverables.
- Bias for iteration: ships rough first versions to learn, then refines smartly. Comfortable with imperfection in early iterations, disciplined about converging.
- Interdisciplinary comfort: Works across computational and experimental boundaries without needing either side translated, and enjoys problems where the data are complex, heterogeneous and dependent on how they were produced.
- Clear communication: explains technical work to engineering, scientific, and non‑technical audiences.
What we offer
- A founding role in Valvian's computational platform, reporting to the Head of Computational Biology & AI.
- Direct collaboration with senior scientific and computational leadership on architectural and scientific direction.
- Competitive salary plus meaningful equity.
- Lisbon‑based hybrid, in‑office presence calibrated to project needs.
- International collaborations and exposure to the European biotech ecosystem.