AI/ML Data Scientist

Medly AI

Greater London

Hybrid

GBP 65,000 - 90,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Medly AI in London is seeking an AI/ML Data Scientist to own how we measure and improve our AI models. You’ll work with founders and engineers who’ve built at Meta and Atlassian, designing benchmarks, evals, and experiments that deliver real improvements for students.

From day one you’ll own the work: develop subject-specific evals, improve offline test suites, run experiments, build tuning datasets, and translate pedagogical judgement into measurable signals.

Qualifications

  • 2+ years in data science, ML, or applied research after graduation.
  • Experience evaluating or benchmarking AI/ML systems with measurable metrics.
  • Strong Python data skills (pandas, numpy) and readable code.
  • Proficient SQL and relational databases.
  • Solid grounding in statistics and experimental design.
  • Familiarity with LLMs, evaluation, prompting, and fine-tuning.
  • Degree in CS, Maths, Physics, Statistics, Engineering, Data Science, or similar.

Responsibilities

  • Own AI benchmarking and create subject-specific evals for tutoring and content generation.
  • Improve eval harnesses and offline test suites to ship model changes confidently.
  • Run experiments and A/B tests on live AI features and interpret results.
  • Build and curate high-quality datasets for fine-tuning and evaluation.
  • Investigate model failures end to end and turn findings into fixes.
  • Collaborate with learning developers and teachers to quantify pedagogical judgement.
  • Work with product data to surface insights and shape future features.
  • Contribute to the codebase (Python, Postgres, AWS, Next.js) where relevant.

Skills

Python
SQL
Statistics
Experiment design
LLMs
Eval frameworks
AWS

Education

Quantitative degree (CS, Maths, Physics, Statistics, Engineering, Data Science)

Tools

AWS

Job description

We're the fastest growing EdTech startup in London, on a mission to change education forever.

This is a moment where AI-native companies are reshaping entire industries, from Cursor in code, to Midjourney in image generation, to Harvey in law. Medly is building that company for education.

Since launch, we've reached over 400,000 students with retention rates that beat Duolingo, and we've delivered personalised learning with real results. We're backed by impact-focused VCs and supported by partners including UCL, Innovate UK, Microsoft, and Google.

The role

We're looking for an AI/ML Data Scientist to join our London team. You'll work directly alongside our founders and engineers who've previously built at Meta, Atlassian, and more, and you'll own how we measure and improve Medly's AI models.

This is the person who tells us the truth about our models. Hundreds of thousands of students rely on our AI to tutor them, mark their work, and generate content pitched at the right level for their exam board. Right now, judging whether a change made that better or worse is the hardest problem we have. You'll build and improve the benchmarks, evals, and experiments that answer it, and you'll be the reason we can ship quickly without breaking what works.

You'll have real ownership from day one. This is not a role where you'll be handed a spec.

What you’ll do
  • - Own our AI benchmarking: help improve and design subject-specific evals that measure tutoring, marking, and content generation quality against what a good teacher would actually say
  • - Improved the eval harnesses and offline test suites that let us ship model and prompt changes with confidence, and catch regressions before students do
  • - Run experiments and A/B tests on live AI features, and make the call on what the results actually mean
  • - Build and curate high-quality datasets for fine-tuning and evaluation, including designing labelling schemes and rubrics that other people can apply consistently
  • - Investigate model failures end to end: find where the AI falls down on notation, working, partial credit, or a specific exam board, and turn that into a fix
  • - Work with our learning developers and teachers to turn pedagogical judgement into something measurable
  • - Get hands-on with our product data: spot patterns, surface insights, and shape what we build next
  • - Contribute to the wider codebase where it makes sense (Python, Postgres, AWS, some Next.js)
What we’re looking for
  • - At least 2 years in a data science, ML, or applied research role, post-graduation
  • - Hands-on experience evaluating or benchmarking AI/ML systems: you've built evals, designed metrics, or run structured model comparisons, and you can talk about where your own metrics misled you
  • - Strong Python for data work (pandas, numpy, notebooks, and the ability to write code other people can run)
  • - Confident with SQL and relational databases
  • - Solid grounding in statistics and experiment design: sampling, significance, confidence intervals, and knowing when a result doesn't mean what it looks like
  • - Working knowledge of LLMs and how they're trained, evaluated, prompted, and fine-tuned
  • - A degree in Computer Science, Maths, Physics, Statistics, Engineering, Data Science, or a similarly quantitative field
  • - Genuine curiosity about how AI systems behave in the wild, not just on a leaderboard
  • - Experience with eval frameworks and LLM-as-judge setups, including their failure modes
  • - Fine‑tuning, RLHF, DPO, or preference data collection
  • - Built annotation or human‑evaluation pipelines, and managed the people doing the labelling
  • - Worked on AI products in production where quality was subjective and hard to measure
  • - Any teaching, tutoring, or exam‑marking experience, or familiarity with the UK curriculum (GCSE, A‑Level, AQA/Edexcel/OCR)
  • - Published research, open‑source contributions, or writing about evaluation
  • - Cloud infrastructure experience (AWS)
  • - Something you've built in your own time that you'd love to show us
Right to work

You must have the right to work in the UK. We're not able to offer visa sponsorship for this role.

Who thrives here
  • - You care about quality and take pride in work that holds up under scrutiny
  • - You're comfortable being the person who says the model got worse, with the evidence to back it
  • - Curious about how things work, and willing to dig into other people's code and data to find out
  • - You take initiative. If something looks off, you flag it or fix it rather than waiting to be told
  • - You can explain a technical result to someone who isn't technical, and be trusted on it
  • - Comfortable in a startup environment. Things move very quickly and priorities shift often
What you’ll get
  • - Real ownership of how Medly measures AI quality, on a product used by hundreds of thousands of students
  • - Direct work with founders and senior engineers with experience at Meta, Atlassian, and beyond
  • - Hybrid setup: 4 days in our central London office, one day work from home
  • - A team that takes your ideas seriously and ships them
  • - Room to grow into ownership of Medly's wider AI and data function
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Junior Fullstack Engineer
Junior Fullstack Engineer

Medly AI • Greater London

On-site
GBP 28,000 - 36,000
AI/ML Data Scientist — Hybrid London, Own AI Benchmarks
AI/ML Data Scientist — Hybrid London, Own AI Benchmarks

Medly AI • Greater London

Hybrid
GBP 65,000 - 90,000
Data Scientist
Data Scientist

Coefficient Systems Ltd • Greater London

On-site
GBP 34,200 - 41,800
33 days paid holiday
Smart Pension contributions
Regular performance reviews
Senior Data Scientist
Senior Data Scientist

VIQU IT • City Of London

On-site
GBP 90,000 - 120,000
Graphics Designer
Graphics Designer

Medly AI • Slough

On-site
GBP 30,000 - 42,000
Mentorship from designers
On-site in central London
Impactful product/brand work
+2
Graphics Designer
Graphics Designer

Medly AI • City Of London

On-site
GBP 26,000 - 38,000
Mentorship from senior designers
On-site in central London office
Brand ownership potential
Engagement Associate
Engagement Associate

Metaview • Greater London

On-site
GBP 70,000 - 110,000
Unlimited AI tools
Equity compensation
High compensation
+2
Lead ML Engineer
Lead ML Engineer

Applied Data Science Partners • Greater London

On-site
GBP 80,000 - 100,000
Market competitive compensation
Annual performance bonus
Generous annual leave
+6
Senior Data Scientist
Senior Data Scientist

VIQU IT • Greater London

On-site
GBP 70,000 - 110,000
Machine Learning Engineering Manager
Machine Learning Engineering Manager

Compare the Market • Peterborough

Hybrid
GBP 110,000 - 140,000