Data Scientist

Kaizo Nederland

Amsterdam

On-site

EUR 55,000 - 90,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Visa sponsorship available
Great office gear
Flexible working schedule
Open holiday policy
Fun workations
Team events

Job summary

Kaizo Nederland is seeking a sharp data scientist to own evaluation and quality loops for the AutoQA product. You can be early in your career, with room to grow into Senior Data Scientist or ML Engineer roles.

The role sits at the intersection of AI and CX, collaborating with enterprise customers and teams in Amsterdam. We offer flexible hours, visa sponsorship for eligible candidates, and a strong growth path in a fast-growing SaaS company.

Qualifications

  • Strong grounding in evaluating AI/ML models and quality metrics.
  • Experience turning customer rubrics into measurable evaluation criteria.
  • Hands-on data analytics with Python and the standard data toolkit.

Responsibilities

  • Translate customer rubrics into AutoQA evaluation instructions.
  • Design and run evaluation experiments on large datasets using LangSmith and BigQuery.
  • Curate data into golden datasets and generate synthetic data for edge cases.
  • Build LLM-as-a-judge pipelines for repeatable evaluation.
  • Make results actionable and collaborate with the AI team to ship improvements.
  • Evaluate retrieval, tool calling, and transcription quality across the stack.

Skills

LLM fundamentals
Python
Pandas
NumPy
Jupyter
SQL

Education

Bachelor's degree in a relevant field

Tools

LangSmith
BigQuery
Docker
Kubernetes

Job description

Do you want to work out how you actually measure whether an AI system is doing a good job, and then make it better? We're looking for a sharp, analytical data scientist to own the evaluation and quality loop of our AutoQA product. You can be early in your career (recent graduates with strong LLM fundamentals are welcome), we have room and a clear path for you to grow.

In a nutshell
  • Join a fast-growing SaaS company in an international environment (steep learning curve guaranteed).

  • Own a high-impact problem: making LLM-powered quality assurance measurably accurate at scale.

  • Sit at the intersection of AI and CX, working directly with enterprise customers and their real‑world QA rubrics.

  • Grow into a Senior Data Scientist, AI Engineer, or ML Engineer role. We invest in progression.

  • Enjoy the perks: flexible hours, open holiday policy, an office in the heart of Amsterdam with hybrid flexibility, visa sponsorship, great gear, workations, and team events.

About Kaizo

At Kaizo, we build a performance development and quality platform for customer support teams. Our AutoQA product uses LLMs to review support conversations against each customer’s own quality rubric, automatically and at scale. Behind it sits a microservices‑based stream processing platform handling over 200 million events per day (Kafka, Kubernetes on Google Cloud, ElasticSearch, MongoDB, BigQuery), and an LLMOps stack built around LangSmith for experimentation, prompt management, and tracing.

The hard part isn’t calling an LLM. It’s knowing, with evidence, how well the system performs on every customer’s unique rubric, and having a reliable, repeatable way to improve it. That’s where you come in.

What you’ll focus on
  • Translate customer rubrics into AutoQA instructions. Work with real customer quality criteria and turn them into precise, testable instructions that LLMs can score reliably.
  • Run experiments that move accuracy. Design and execute evaluation experiments on large, representative datasets using LangSmith and BigQuery, and track quality with our performance metrics.
  • Build the datasets that make evaluation possible. Curate raw production data into golden datasets with balanced coverage, and generate synthetic data to cover the rare cases that matter most. QA is a discipline of rare occurrences: distributions are skewed, positives are scarce, and resourcefulness beats volume.
  • Build LLM‑as‑a‑judge pipelines to assess system quality internally and make evaluation repeatable.
  • Make results actionable. Your experiments should end in a decision: change a prompt, adjust which tools the system uses, surface context the AI is missing, or flag where new capabilities are needed. You’ll work with the AI team to ship those decisions.
  • Evaluate across the full stack. Beyond scoring quality, you’ll help validate retrieval (RAG/IR), tool calling, and speech pipelines (transcription and diarization quality).
  • Join customer calls with the team to understand how QA leaders define quality, and feed what you learn back into the product.
What you’ll grow into
  • Shaping how customers monitor quality themselves, catch drift, and keep their AutoQA setup improving over time.
  • Smarter categorization and routing of conversations to power analytics and get the right tickets to the right evaluation.
  • A solid understanding of how LLMs work and hands‑on experience prompting them for accuracy (coursework, thesis, internships, or side projects all count; production experience is a bonus).
  • Good applied statistics: experiment design, classifier evaluation, precision/recall trade‑offs, and working with heavily imbalanced data.
  • Strong Python skills and fluency with the standard data toolkit (Pandas, NumPy, Jupyter). SQL is a plus.
  • An analytical, evidence‑first mindset: you’d rather measure than assume.
  • Product sense and empathy for end users. You’ll be building for QA managers and support agents, not just for benchmarks.
  • Excellent written and verbal communication. You’ll present findings to the team and join customer conversations.
  • 0 to 2 years of industry experience. Recent graduates with strong relevant work are encouraged to apply.
  • A team player who’s comfortable wearing multiple hats. We’re an early‑stage company and things move fast.

Bonus points for:

  • Experience with LangSmith or similar LLMOps/evaluation tooling
  • Google Cloud Platform, BigQuery, Docker, or Kubernetes
  • Speech/audio processing or ASR evaluation
  • Fine‑tuning or building synthetic datasets for LLMs
Who you’ll work with

You’ll join our AI team (two data scientists, an AI engineer, and an ML engineer) and collaborate closely with our data engineers, frontend engineers, designer, product manager, and CX teams. You’ll have mentorship from day one and real ownership fast.

What’s in it for you?
  • An office in the heart of Amsterdam, with the flexibility of hybrid working
  • Visa sponsorship available for eligible candidates
  • Great office gear: MacBook, tools, desk, chair, whatever you need
  • Flexible working schedule and an open holiday policy
  • Fun workations and team events
  • A clear growth path into Senior Data Scientist, AI Engineer, or ML Engineer roles
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Slashhash • Amsterdam

Hybrid
EUR 60,000 - 90,000
Data Scientist: LLM QA Evaluation & Metrics
Data Scientist: LLM QA Evaluation & Metrics

Kaizo Nederland • Amsterdam

Hybrid
EUR 55,000 - 90,000
Visa sponsorship available
Great office gear
Flexible working schedule
+3
Data Scientist — AutoQA & LLM Evaluation
Data Scientist — AutoQA & LLM Evaluation

Slashhash • Amsterdam

Hybrid
EUR 60,000 - 90,000
AI Engineer (Hybrid - Zaandam, Netherlands) Data Analytics & Artificial Intelligence · Netherlands ·
AI Engineer (Hybrid - Zaandam, Netherlands) Data Analytics & Artificial Intelligence · Netherlands ·

Keyrus Portugal • Zaandam

Hybrid
EUR 70,000 - 89,000
Meal allowance
Private medical insurance
Annual leave 22-25 days
+2
Data Scientist (LLM)
Data Scientist (LLM)

Qogita • Amsterdam

Hybrid
EUR 60,000 - 75,000
Annual leave
Bonus program
Equity package
+6
Senior AI Engineer
Senior AI Engineer

allyourbi • Rotterdam

Hybrid
EUR 81,000 - 91,000
Growth budget €5k/yr
25 vacation days
Buy up to 10 days
+5
(MSc/PhD) AI Research Intern - LLMs, Causal Inference & Decision Making
(MSc/PhD) AI Research Intern - LLMs, Causal Inference & Decision Making

Prosus • Amsterdam

On-site
EUR 13,000 - 23,000
Learn from the best
Hard ML problems
Unique data at scale
+5
AI / ML Engineer
AI / ML Engineer

Awesome Compliance Technology BV • Amsterdam

Hybrid
EUR 90,000 - 130,000
Early team role
Equity
Culture of shipping
+2
Data ML Engineering Lead
Data ML Engineering Lead

kaiko.ai • Amsterdam

Hybrid
EUR 80,000 - 120,000
Competitive salary
Good pension plan
25 vacation days per year
+2
Founding AI Engineer
Founding AI Engineer

Visa Hunt • Netherlands

On-site
EUR 130,000 - 165,000
Stock options
On-site in Amsterdam
No visa sponsorship