ML Research Engineer, Evaluation

PassFort

Greater London

On-site

GBP 65,000 - 100,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Novogaia is seeking a machine learning research engineer to design and run rigorous benchmarks for molecular AI systems. You will own evaluation pipelines, construct adversarial tests, and translate results into clear research priorities for the team.

You will collaborate with AI engineers, computational biologists, and chemists to ensure benchmarks reflect real discovery challenges and provide defensible conclusions about model utility.

Qualifications

  • Experience building or evaluating ML benchmarks.
  • Excellent technical communication to researchers and non-technical stakeholders.
  • Ability to analyze model behavior critically.
  • Strong Python proficiency and containerized pipelines.
  • Familiarity with cheminformatics formats (SMILES, InChIKey).
  • Interest in computational mass spectrometry.

Responsibilities

  • Develop a deep understanding of Novogaia's models, data, and evaluation needs.
  • Design benchmark tasks reflecting real discovery problems across multiple modalities.
  • Build datasets and controls ensuring benchmark trustworthiness: hard negatives and leakage-safe splits.
  • Communicate evaluation results clearly to the modeling team and document findings.
  • Lead evaluation work, define what good looks like, and translate results into priorities.

Skills

Benchmark design
Python programming
Model evaluation
Scientific communication
Mass spectrometry familiarity
Cheminformatics basics

Tools

Docker
Git
MLflow

Job description

Overview

Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.


We are a small team of AI engineers, computational biologists and chemists building foundation models for molecular structure prediction from mass spectrometry data. We are seeking a machine learning research engineer who excels at designing rigorous tests for molecular AI systems, and who can turn \"does this model actually work\" into a concrete, defensible answer. You will own the design of our internal benchmarks and evaluation pipelines; build the adversarial checks that catch shortcut learning and leakage. You will work closely with our modeling team to translate evaluation results into research priorities.


The Role


  • Develop a deep understanding of Novogaia's models, data, and evaluation needs


  • Design benchmark tasks that reflect real discovery problems: de novo molecular structure generation, molecular and spectral retrieval, mass spectrum simulation, molecular formula prediction, and compound property/class prediction


  • Build the datasets and controls that make a benchmark trustworthy: hard negatives, leakage-safe splits, and null baselines that catch a model exploiting shortcuts instead of genuine signal


  • Communicate evaluation results as clear findings for the modeling team, and as documentation and data cards that others can trust and reproduce


  • Work across the team to scope and lead evaluation work, including:



    • Defining what \"good\" looks like for a given model or task, and choosing the right test for it


    • Analyzing model behavior and interpreting results for researchers and non-technical stakeholders alike


    • Working with engineers to turn one-off analyses into repeatable, reproducible evaluation pipelines




  • Translate lessons from evaluation work into research priorities, data requirements, and R&D direction



In your first year, you'll take the lead on how Novogaia measures model quality, from benchmark design through adversarial testing and reporting. The tests you build will decide how much weight anyone can put on our models' outputs.


What We Require


  • Research or applied experience in machine learning, with direct experience building or rigorously evaluating ML benchmarks


  • Exceptional technical communication skills, including the ability to explain evaluation findings clearly to both researchers and non-technical stakeholders


  • Ability to analyze model behavior and interpret computational results critically


  • Strong proficiency in Python, and comfort with reproducible, containerized pipelines


  • Familiarity with cheminformatics representations (SMILES, InChIKey, molecular fingerprints), or willingness to pick these up quickly


  • Familiarity with computational mass spectrometry or eagerness to learn



What We Value


  • Ability to dive deep into a result until you know whether it's real or an artifact


  • Strong scientific judgment and a willingness to question the benchmark's own assumptions, not just the model's


  • Motivation to build infrastructure other people can confidently rely and build upon


  • Comfortable being the person who tells the team a result doesn't hold up


  • Curiosity, low ego, and a willingness to learn the chemistry side quickly, even if it's outside your original training


Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Benchmark Architect for Drug Discovery
ML Benchmark Architect for Drug Discovery

PassFort • Greater London

On-site
GBP 65,000 - 100,000
Senior / Staff Machine Learning Scientist
Senior / Staff Machine Learning Scientist

your Jared • Glasgow

On-site
GBP 90,000 - 150,000
Research Engineer - Post-Training
Research Engineer - Post-Training

Helical • Greater London

On-site
GBP 110,000 - 150,000
ML Research Engineer, London London
ML Research Engineer, London London

Isomorphic Labs Limited • Greater London

On-site
GBP 60,000 - 80,000
ML Research Engineer
ML Research Engineer

Hlx Life Sciences • United Kingdom

On-site
GBP 60,000 - 90,000
Competitive salary
Meaningful equity
Support for conferences and publications
Research Engineer, Evals - Member of Technical Staff
Research Engineer, Evals - Member of Technical Staff

Callosum Technologies Ltd. • Greater London

On-site
GBP 90,000 - 140,000
Competitive salary
Equity & Ownership
Private healthcare
+2
Research Engineer, Evals - Member of Technical Staff
Research Engineer, Evals - Member of Technical Staff

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Equity ownership
Private healthcare
+1
Applied ML Engineer/Scientist
Applied ML Engineer/Scientist

Boltz • Greater London

On-site
GBP 70,000 - 110,000
Equity ownership
ML Research Engineer, London
ML Research Engineer, London

Isomorphic Labs • Greater London

On-site
GBP 65,000 - 95,000
Research Engineer (LLM Performance), London
Research Engineer (LLM Performance), London

isomorphiclabs • Greater London

On-site
GBP 90,000 - 150,000