Senior Data Scientist — AI Evaluation & Quality | Internal AI Agents (Remote in Europe)

Lever, Inc.

Barcelona

Híbrido

EUR 70.000 - 110.000

Jornada completa

14 días+
Generador de candidaturas

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Stock options
Work & Swim program

Descripción de la vacante

Finom’s AI Team seeks a seasoned data scientist to own evaluation methodologies for its internal AI agents. You will design benchmarks, set quality gates, and partner with domain experts to label cases and resolve disagreements.

You’ll work with Databricks, Claude Code, and dbt to build robust evaluation pipelines. You will lead the offline and online evaluation suites across ~20 processes and translate findings into actionable product decisions, balancing quality, cost, and latency.

Formación

  • 5+ years in Data Science / Product Analytics / Applied AI roles with product-level metric ownership.
  • Production LLM Experience: built, shipped, or evaluated LLM-based systems as core responsibility in last 1–2 years.
  • Autonomous Quality Ownership: owned evaluation methodology or analytics for an entire product or end-to-end process.
  • Fluent Python & SQL: write clean data pipelines, evaluation harnesses, and dbt transformations.
  • Statistical Rigor: sampling, hypothesis testing, variance analysis, and confidence intervals on noisy metrics.

Responsabilidades

  • Own and extend offline/online evaluation suites across ~20 internal AI agent processes.
  • Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines.
  • Label cases with domain experts and resolve annotator disagreements.
  • Build test datasets from real user and operational traffic.
  • Harden statistical methodology and measure metric shifts vs. noise.
  • Translate quality metrics into operational decisions with process owners.

Conocimientos

Data Science
Product Analytics
Applied AI
Python
SQL
Statistical Rigor

Herramientas

Databricks
DeepEval
Claude Code
Cursor
Python
SQL
dbt

Descripción del empleo

About Finom

Finom is a European tech startup headquartered in Amsterdam, and we’re on a journey towards revolutionizing the financial landscape for entrepreneurs worldwide. Our mission is to develop an all-in-one financial B2B solution that integrates banking functions, accounting, financial management, and invoicing into a seamless, mobile-first platform.

We recently closed a €115 million Series C equity round (around $133 million), bringing our total funding to approximately $346 million. This significant investment follows a $105 million growth funding round from General Catalyst, a long-term backer since 2021 known for supporting companies like Airbnb, HubSpot, KAYAK, and Stripe.

Finom's platform goes beyond traditional banking, offering invoicing and a growing suite of features, including AI-enabled accounting, aiming to simplify financial management for entrepreneurs. We're actively expanding our reach across key EU markets like Germany, France, the Netherlands, Italy, and Spain.

At Finom, we’re not just redefining the entrepreneurial experience — we’re empowering our employees to make a real difference. Your work matters, and your impact extends far beyond product metrics. We nurture innovation and an inspiring work environment where bold ideas thrive, prioritizing thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and our business as a whole.

Maintaining our start-up spirit, we prioritize thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and, of course, our business.

AI Team

You’ll join Finom’s AI Team as the founding IC dedicated to the quality, evaluation and telemetry of AI agents powering Finom's internal operations and tools (including Ops workflows and internal AI analytics engines across ~20 core processes).

Our belief

An AI agent is only as good as the evaluation loop running on it. Because internal operational agents directly touch financial, compliance and support workflows, evaluation is an exact statistical and engineering discipline here.

Your mission

Design the evaluation methodology, build golden benchmarks, and establish quality gates for our internal AI agents from scratch—working directly with process owners and domain experts, with no senior quality owner above you to lean on.

Core Stack

Databricks, DeepEval, Claude Code, Cursor, Python, SQL, dbt.

What You Will Be Doing
  • Own and extend our offline/online evaluation suites across ~20 internal AI agent processes—datasets (capability + regression), LLM-as-a-judge rubrics, and deterministic checks.
  • Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines before agent prompt, context, or tool changes hit production.
  • Work directly with domain experts to label cases and resolve annotator disagreement—fixing definition criteria rather than averaging disagreement away.
  • Build test datasets derived from real user & operational traffic (tickets, internal chats, colleague queries) rather than synthetic edge cases.
  • Harden statistical methodology: handle judge drift, verbosity bias, non-determinism, and measure true metric shifts vs. noise.
  • Translate quality numbers into operational decisions: run weekly syncs with process owners to define clear quality vs. cost/latency trade-offs.
Must-Haves
  • 5+ years in Data Science / Product Analytics / Applied AI roles, with sustained product-level metric ownership.
  • Production LLM Experience: In the last 1–2 years, you have built, shipped, or evaluated LLM-based systems (RAG, multi-step tool use, agents) as a core, primary job responsibility.
  • Autonomous Quality Ownership: Proven track record of owning evaluation methodology or analytics for an entire product or end-to-end process (what to build vs. what NOT to build).
  • Fluent Python & SQL: Ability to write clean data pipelines, evaluation harnesses, and dbt transformation models directly.
  • Statistical Rigor: Applied knowledge of sampling, hypothesis testing, variance analysis, and confidence intervals on noisy metrics.
Daily Setup & Tooling
  • AI-assisted coding (Claude Code, Cursor, or Codex) is your default daily authoring environment for Python, SQL, and evaluation scripts—not something you occasionally experiment with.
  • You can walk us through concrete work tasks from the last month where AI coding tools accelerated your engineering and data analysis.
How we work — one thing we mean seriously
  • AI-assisted coding is our default authoring environment, not a bonus
  • Claude Code is our main tool — you'll reach for it for SQL, Python, analyses, dashboards, and internal scripts
  • We're looking for analysts who are already curious and fluent with AI coding — or genuinely excited to become fluent fast
  • We care about what you ship and how clearly you think
  • If this idea excites you rather than worries you, you'll feel at home here
What You Will Get In Return
  • Make a genuine impact on the product Join our upward trajectory, and grow with us. We provide the resources and opportunities for continuous personal and professional development, empowering you to make a genuine impact on our evolving product.
  • Work in the EU Embark on this exciting journey with us and enjoy the flexibility of traveling and working remotely or in a hybrid model across Europe.
  • Become a stock options holder Unlock your inner entrepreneur and align your aspirations with ours through our Stock Options Program. This exciting opportunity is available to every team member, from junior team members to our founders.
  • Receive unwavering support and care Finom stands by you at every step, embodying our commitment to your well-being and success reflected in our modern, friendly, and eco-conscious corporate culture. We offer constant support and care to ensure your Finom experience is successful and fulfilling.
  • Work & Swim program Immerse yourself in our exclusive Work & Swim Program. Spend one month in a comfortable corporate apartment in enchanting Cyprus. It's the ideal opportunity to strike the perfect work-life balance while enjoying breathtaking Mediterranean views.
Equal Opportunity Statement

At Finom, we're an equal opportunity employer and value diversity at our company. We embrace diversity and invite applications from all walks of life. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, disability status, or other applicable legally protected characteristics.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Data Scientist — AI Evaluation & Quality (Remote)
Senior Data Scientist — AI Evaluation & Quality (Remote)

Finom • Barcelona

Presencial
EUR 45.000 - 75.000
Stock options
Remote/Hybrid across Europe
Work & Swim program
+1
Senior AI Engineer (Remote)
Senior AI Engineer (Remote)

Finom • Barcelona

Presencial
EUR 110.000 - 140.000
Stock options
Work & Swim program
Remote/hybrid across Europe
Lead AI Engineer (Remote)
Lead AI Engineer (Remote)

Finom • Barcelona

A distancia
EUR 120.000 - 180.000
Stock Options Program
Work & Swim Program
Equal Opportunity Policy
Product Analytics Chapter Lead
Product Analytics Chapter Lead

Finom • España

A distancia
EUR 60.000 - 90.000
Stock options program
Work & Swim program (Cyprus)
Remote work across Europe
People Product Specialist
People Product Specialist

Finom • Barcelona

Presencial
EUR 42.000 - 64.000
Stock options
Hybrid work across Europe
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Finom • España

Híbrido
EUR 90.000 - 130.000
Stock options
Hybrid/Remote across Europe
Work & Swim Cyprus program
+1
Principal Product Manager - Finom AI
Principal Product Manager - Finom AI

DutchTechX • España

A distancia
PHP 6.508.000 - 9.400.000
Stock options
Remote-friendly across Europe
Work & Swim Program
+1
Senior Product Manager
Senior Product Manager

Finom • España

Presencial
EUR 90.000 - 125.000
Stock options
Work & Swim Program
Remote work across Europe
Credit Risk Manager (SME Lending)
Credit Risk Manager (SME Lending)

Finom • España

Híbrido
EUR 110.000 - 160.000
Stock options
Work & Swim program in Cyprus
Remote/hybrid across Europe
Enterprise Risk & Reporting Lead
Enterprise Risk & Reporting Lead

Finom • España

Presencial
EUR 90.000 - 140.000
Stock options
Work & Swim Cyprus program
Career development