Data engineer

Cato

Barcelona

Presencial

EUR 40.000 - 60.000

Jornada completa

Hace 7 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Cato is hiring a data engineer to build and run the pipelines that move tenders from source portals to customers. You own ingestion, merge, enrichment, and delivery end to end, with real ownership over data sources, jobs and tables.

You will work on scrapers for portals in Spain and Italy, write robust orchestrations (Prefect, Airflow, Dagster, Argo), perform AI enrichment with LLM extraction, and ensure production SQL performance.

Formación

  • 2–4 years building data pipelines in production.
  • Hands-on with PostgreSQL beyond writing queries—you've had to make one fast.
  • Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.

Responsabilidades

  • Own scrapers and ingestion for a set of national portals across Spain and Italy.
  • Write and maintain orchestrator flows: retries, backfills, alerting, and validation.
  • Work on merge, dedup and reconciliation to deliver a single correct tender.
  • Ship AI enrichment steps: batch LLM extraction, embeddings, OCR on attachments.
  • Write production SQL with query plans, indexes, JSONB, partitioning, and migrations.
  • Guard data quality with tests and checks that fail loudly.

Conocimientos

SQL mastery
Python production
Data pipelines
Automation mindset
Troubleshooting

Herramientas

PostgreSQL
Prefect
Airflow
Dagster
ArgoWorkflows
JSONB

Descripción del empleo

Build and run the pipelines that carry a tender from the source portal to the customer's screen: ingestion, merge, enrichment, delivery. You own concrete pieces of the data pipeline end to end — not tickets handed to you, but the sources, jobs and tables behind them.

What you'll actually do
  • Own scrapers and ingestion for a set of national portals across Spain and Italy, where reading the source in its original language is part of the job.
  • Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?"
  • Work on merge, dedup and reconciliation — the same tender arrives three times, in three shapes, and only one version can reach the customer.
  • Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments.
  • Write SQL that survives production: query plans, indexes, JSONB, partitioning, CONCURRENTLY migrations.
  • Guard data quality with tests and checks that fail loudly before a customer finds the gap.
Ideal profile
  • Real SQL: you can read an EXPLAIN and say why the plan is bad, not just that it is slow.
  • Python you'd put in production: typed, tested, and readable six months later.
  • Pipelines you've actually operated: with an orchestrator (Prefect, Airflow, Dagster, ArgoWorkflows) and the 3 a.m. failures that come with them.
  • Builder by default: you see a manual process and your first instinct is to automate it.
  • Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans.
Experience
  • 2–4 years building data pipelines in production.
  • Hands-on with PostgreSQL beyond writing queries — you've had to make one fast.
  • Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.
What you won't find here
  • No micromanagement: we trust you to own your part of the stack.
  • No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile.
  • No "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns.
  • AI: batch LLM extraction, embeddings, OCR
Compensation

RAL €40,000 – €60,000 + equity, depending on seniority and profile.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Junior data engineer - España
Junior data engineer - España

Cato • Barcelona

Presencial
EUR 45.000 - 65.000
Data Engineer - End-to-End Pipelines & AI Enrichment
Data Engineer - End-to-End Pipelines & AI Enrichment

Cato • Barcelona

Presencial
EUR 40.000 - 60.000
Data Engineer (Madrid based)
Data Engineer (Madrid based)

Auctane • Madrid

Presencial
EUR 43.000 - 52.000
Salary range 43,000–52,000 € per year
Cobee benefits for expenses
Private health insurance via Cigna
+13
Data Engineer I, Data Technology And Products
Data Engineer I, Data Technology And Products

Amazon Spain Services, S.L.U. • Barcelona

Presencial
EUR 65.000 - 95.000
Data Engineer
Data Engineer

Senovo IT Ltd • Madrid

Híbrido
EUR 40.000 - 60.000
Data Engineer
Data Engineer

Amazon • Barcelona

Presencial
EUR 65.000 - 90.000
Data Engineer
Data Engineer

DEUS: human(ity)-centered AI • Galicia

Presencial
EUR 35.000 - 60.000
Top hardware and software provided
Flexible working hours
Inclusive workplace culture
+1
Data Software Engineer
Data Software Engineer

The Workshop • Madrid

Híbrido
EUR 45.000 - 65.000
Private life and health insurance
Pension plan
Gym reimbursement
+2
Data Engineer - German speaker
Data Engineer - German speaker

isolutions AG • España

Híbrido
EUR 55.000 - 85.000
Fringe benefits package
Hybrid or remote in Spain
Training budget
+3
Data Analitics - Industrial Digital Platform
Data Analitics - Industrial Digital Platform

Verdalia Bioenergy • Madrid

Híbrido
EUR 50.000 - 70.000
Strategic role with real impact
Dynamic and fast-growing environment
Opportunity to work on a modern data platform