Junior data engineer - España

Cato

Barcelona

Presencial

EUR 45.000 - 65.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Destaca para este puesto: genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Descripción de la vacante

Cato in Barcelona seeks a Python-focused data engineer to own tender-scraping end-to-end, from source to customer-ready data. You will build and maintain scrapers for national portals, handle HTML/XML content, manage pagination and sessions, and implement AI enrichment steps.

Ideal candidates are proficient in Python and SQL, comfortable with messy sources, and able to automate manual steps. You will work with orchestration tools and ensure reliable, auditable data delivery.

Formación

  • 1–2 years writing Python in production for scrapers or ETL.
  • Experience with HTTP/HTML/XML parsing and pagination.
  • Comfort with messy sources and data quality checks.

Responsabilidades

  • Build and maintain scrapers for national tender portals.
  • Turn messy sources into clean records and ensure reliability.
  • Write orchestrator flows with retries and alerts.
  • Work on merge/deduplication and AI enrichment steps.
  • Guard data quality with tests and checks before delivery.

Conocimientos

Python (typed, tested)
HTTP parsing
HTML parsing
XML parsing
Pagination strategies
SQL (joins, windows)
Automation mindset
LLM extraction familiarity

Herramientas

Prefect
Airflow
Dagster
OCR on attachments

Descripción del empleo

Bring a tender from the source portal into Cato: scraping, parsing, merging, enrichment. You'll start by owning a handful of sources end to end - the scraper, the job behind it, and the data that comes out - and take on more as you go. Not tickets handed to you: sources you're responsible for.

What you'll actually do
  • Build and maintain scrapers for national tender portals, where reading the source in its original language is part of the job.
  • Keep them alive: portals change their HTML, move endpoints, break pagination, throttle you. You find out before the customer does.
  • Turn messy sources into clean records: broken HTML, inconsistent XML, APIs that lie about their own schema.
  • Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?"
  • Work on merge and dedup - the same tender arrives three times, in three shapes, and only one version can reach the customer.
  • Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments.
  • Guard data quality with tests and checks that fail loudly before a customer finds the gap.
Ideal profile
  • Python that holds up: typed, tested, and readable six months later.
  • You've scraped something real: HTTP, HTML and XML parsing, pagination, sessions, rate limits - and you know why a scraper that worked yesterday is broken this morning.
  • SQL you're comfortable in: joins, aggregations, window functions. You'll read from the database every day; tuning and running it isn't your job.
  • Builder by default: you see a manual process and your first instinct is to automate it.
  • Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans.
  • You close your own loop: you check that what you shipped actually ran, before someone else has to ask.
Experience
  • 1-2 years writing Python in production: scrapers, ETL scripts, automation - anything that had to run unattended and be fixed when it didn't.
  • Exposure to an orchestrator (Prefect, Airflow, Dagster) is a plus, not a requirement: you'll learn ours properly.
  • Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.
What you won't find here
  • No micromanagement: we trust you to own your part of the stack.
  • No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile.
  • No "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns.
  • AI: batch LLM extraction, embeddings, OCR
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Data engineer
Data engineer

Cato • Barcelona

Presencial
EUR 40.000 - 60.000
Junior Data Engineer: Scrapers, ETL & AI Enrichment
Junior Data Engineer: Scrapers, ETL & AI Enrichment

Cato • Barcelona

Presencial
EUR 45.000 - 65.000
Data Engineer - End-to-End Pipelines & AI Enrichment
Data Engineer - End-to-End Pipelines & AI Enrichment

Cato • Barcelona

Presencial
EUR 40.000 - 60.000
Data Engineer (Madrid based)
Data Engineer (Madrid based)

Auctane • Madrid

Presencial
EUR 43.000 - 52.000
Salary range 43,000–52,000 € per year
Cobee benefits for expenses
Private health insurance via Cigna
+13
Data Software Engineer
Data Software Engineer

The Workshop • Madrid

Híbrido
EUR 45.000 - 65.000
Private life and health insurance
Pension plan
Gym reimbursement
+2
Data Crawler Specialist
Data Crawler Specialist

Delectatech • Barcelona

Presencial
EUR 30.000 - 35.000
Contrato indefinido
Horario flexible
Oficina en Barcelona
+1
Senior Data Engineer
Senior Data Engineer

Sabia Personal • Valencia

Presencial
EUR 60.000 - 90.000
Contrato indefinido
Igualdad de salario garantizada
Bono de rendimiento
+5
Junior Data Engineer
Junior Data Engineer

Smadex • Barcelona

Híbrido
EUR 30.000 - 45.000
Great compensation package
Meal vouchers
Monthly gym allowance
+2
AI Engineer (Full-Stack)
AI Engineer (Full-Stack)

Bifrost Studios • Barcelona

Presencial
Confidential
Security and GDPR compliance
Supabase in production
Multi-tenant SaaS experience
+1
Data Engineer
Data Engineer

Senovo IT Ltd • Madrid

Híbrido
EUR 40.000 - 60.000