Junior data engineer - Italia

Cato

Milano

In loco

EUR 35.000 - 45.000

Tempo pieno

10 giorni fa
Generatore di candidature

Non inviare un curriculum generico — genera un curriculum e una lettera di presentazione personalizzati per questo specifico ruolo.

Supera i filtri ATS

Descrizione del lavoro

Cato is seeking a Python developer to own and operate scrapers for national tender portals, handling the full data lifecycle from extraction to enrichment. You will maintain the scraper, job behind it, and the data output, ensuring reliability even as sources change.

You’ll work with SQL, write robust ETL scripts, and leverage orchestrators like Prefect, Airflow, or Dagster. Strong curiosity about AI enrichment and data quality is essential.

Competenze

  • 1–2 years of Python in production (scrapers, ETL scripts, automation).
  • Experience reading from databases daily; tuning and running queries is common.
  • Exposure to orchestrators (Prefect, Airflow, Dagster) is a plus, not required.

Mansioni

  • Build and maintain scrapers for national tender portals and end-to-end ownership of sources.
  • Keep scrapers alive as portals change HTML, endpoints, and pagination.
  • Turn messy sources into clean records: broken HTML, inconsistent XML, APIs with schema issues.
  • Write and maintain orchestrator flows: retries, backfills, alerting, and run verification.

Conoscenze

Python
SQL
ETL
Web scraping
Automation
Orchestrators

Strumenti

Prefect
Airflow
Dagster
OCR

Descrizione del lavoro

Your mission

Bring a tender from the source portal into Cato: scraping, parsing, merging, enrichment. You'll start by owning a handful of sources end to end - the scraper, the job behind it, and the data that comes out - and take on more as you go. Not tickets handed to you: sources you're responsible for.

What you'll actually do

Build and maintain scrapers for national tender portals, where reading the source in its original language is part of the job. Keep them alive: portals change their HTML, move endpoints, break pagination, throttle you. You find out before the customer does. Turn messy sources into clean records: broken HTML, inconsistent XML, APIs that lie about their own schema. Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to "Did today's run actually land?" Work on merge and dedup - the same tender arrives three times, in three shapes, and only one version can reach the customer. Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments. Guard data quality with tests and checks that fail loudly before a customer finds the gap.

Ideal profile

Python that holds up: typed, tested, and readable six months later. You've scraped something real: HTTP, HTML and XML parsing, pagination, sessions, rate limits - and you know why a scraper that worked yesterday is broken this morning. SQL you're comfortable in: joins, aggregations, window functions. You'll read from the database every day; tuning and running it isn't your job. Builder by default: you see a manual process and your first instinct is to automate it. Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans. You close your own loop: you check that what you shipped actually ran, before someone else has to ask. Experience 1-2 years writing Python in production: scrapers, ETL scripts, automation - anything that had to run unattended and be fixed when it didn't. Exposure to an orchestrator (Prefect, Airflow, Dagster) is a plus, not a requirement: you'll learn ours properly. Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.

What you won't find here

No micromanagement: we trust you to own your part of the stack. No "standard" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile. No "we've always done it this way" excuses: we're here to disrupt, not to follow old patterns.

Our Tech Stack

Data & Infra: Python, PostgreSQL, Prefect, AWS AI: batch LLM extraction, embeddings, OCR

Compensation

RAL €35,000 - €45,000 equity, depending on profile.

Hiring Manager

Lorenzo Rossetto lrossetto@get-cato.com

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Data Engineer II (Onsite – Rome)
Data Engineer II (Onsite – Rome)

Agile Lab S.r.l. • Roma

In loco
EUR 38.000 - 49.000
Training monthly budget
Structured career path
Relocation support
+3
Data Engineer I (On Site - Rome)
Data Engineer I (On Site - Rome)

Agile Lab S.r.l. • Roma

In loco
EUR 30.000 - 40.000
Training budget €1,000/year
Salary progression on ladder
Conference attendance support
+5
Data Engineer
Data Engineer

Cortilia • Milano

Ibrido
EUR 35.000 - 55.000
Piano welfare aziendale
Smart working in alcuni giorni della settimana
Buoni pasto
+3
Data Engineer - freelance 12 months (AI Team)
Data Engineer - freelance 12 months (AI Team)

Moneyfarm • Cagliari

In loco
EUR 60.000 - 90.000
Product Engineer, Platform
Product Engineer, Platform

Callimacus • Milano

In loco
EUR 45.000 - 65.000
Join Us: Data Engineering To Unlock Value From Data (Milano-Torino)
Join Us: Data Engineering To Unlock Value From Data (Milano-Torino)

Target Reply • Trentino-Alto Adige

In loco
EUR 70.000 - 110.000
Senior Data Scientist - Digital Pills
Senior Data Scientist - Digital Pills

Welyk • Torino

In loco
EUR 55.000 - 65.000
13th month salary
Annual bonus
AI Engineer
AI Engineer

TeamSystem S.p.A. • Milano

Ibrido
EUR 34.000 - 40.000
Wellbeing Digital Wallet
Light Friday
Opportunità di crescita
Product Engineer, Agents
Product Engineer, Agents

Callimacus • Milano

In loco
EUR 45.000 - 65.000
Growth opportunities
Equity
Milan offices
+3
Data Engineer I
Data Engineer I

Docebo • Bardi

Ibrido
EUR 29.000 - 39.000
Employee Share Purchase Plan 15%
Health benefits
Paid vacation days
+5