Data Scrapping Engineer

Trabajosihay

Colombia

Presencial

COP 66.960.000 - 133.920.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Trabajosihay is seeking a Data Scraper to collect, organize, and normalize data from public and government sources into a consistent, structured format. You will work on data acquisition challenges, investigate unfamiliar sources, and transform outputs into predefined formats for downstream systems.

The role emphasizes handling messy datasets, researching source structures, and creating reusable automation workflows with strong documentation and independent work styles.

Formación

  • Strong experience with web scraping and data extraction.
  • Practical programming experience using Python.
  • Experience with HTML parsing, APIs, HTTP requests, FTP sources, and structured or unstructured data.
  • Ability to evaluate, debug, and improve scraping solutions.
  • Strong analytical and problem-solving skills.
  • Experience building reusable automation workflows rather than one-off scripts.
  • Familiarity with relational databases (PostgreSQL preferred) and a normal Git workflow.
  • Strong documentation and communication skills.
  • Ability to work independently and take ownership of technical challenges.
  • High attention to detail and commitment to data accuracy.

Responsabilidades

  • Research and identify public and government data sources.
  • Extract and normalize data from websites, APIs, feeds, and online repositories.
  • Build reusable, maintainable, and re-runnable scripts and scraping workflows.
  • Deliver structured outputs in predefined formats.
  • Provide sample outputs for review before processing larger datasets.
  • Document data sources, extraction methodologies, challenges encountered, and re-run procedures.
  • Capture and report any relevant information discovered during extraction, including inconsistencies, amendments, effective dates, repeal notes, or related metadata.
  • Troubleshoot data acquisition issues and propose alternative approaches when needed.
  • Collaborate with stakeholders through regular check-ins and written communication.
  • Maintain version-controlled code repositories and follow standard development practices.

Conocimientos

Web Scraping
Data Extraction
Python
HTML Parsing
APIs
Problem Solving
Data Normalization
Data Transformation
Git
Structured Data
Unstructured Data
Data Validation
Public Data Sources
Government Data
Technical Documentation
Playwright
Selenium
Puppeteer
Scrapy
PostgreSQL

Herramientas

PostgreSQL
Git
Playwright
Selenium
Puppeteer
Scrapy

Descripción del empleo

Job Description:

We are seeking a Data Scraping to help collect, organize, and normalize data from public and government sources into a consistent, structured format.

This role focuses on solving complex data acquisition challenges, researching unfamiliar sources, extracting information from websites and feeds, and transforming it into predefined formats that can be consumed by downstream systems.

The ideal candidate enjoys working with messy datasets, investigating how websites and data sources are structured, and creating reusable solutions that can be executed repeatedly with consistent results.

This position requires strong problem‑solving skills, attention to detail, and the ability to work independently while documenting findings and processes clearly.

Schedule:

Monday to Friday – 12:00 PM – 8:00 PM CST

Responsibilities:
  • Research and identify public and government data sources.
  • Extract and normalize data from websites, APIs, feeds, and online repositories.
  • Build reusable, maintainable, and re‑runnable scripts and scraping workflows.
  • Deliver structured outputs in predefined formats.
  • Provide sample outputs for review before processing larger datasets.
  • Document data sources, extraction methodologies, challenges encountered, and re‑run procedures.
  • Capture and report any relevant information discovered during extraction, including inconsistencies, amendments, effective dates, repeal notes, or related metadata.
  • Troubleshoot data acquisition issues and propose alternative approaches when needed.
  • Collaborate with stakeholders through regular check‑ins and written communication.
  • Maintain version‑controlled code repositories and follow standard development practices.
Qualifications:
  • Strong experience with web scraping and data extraction.
  • Practical programming experience using Python or similar scripting languages.
  • Experience working with HTML parsing, APIs, HTTP requests, FTP sources, and structured or unstructured data.
  • Ability to evaluate, debug, and improve scraping solutions.
  • Strong analytical and problem‑solving skills.
  • Experience building reusable automation workflows rather than one‑off scripts.
  • Familiarity with relational databases (PostgreSQL preferred) and a normal Git workflow.
  • Strong documentation and communication skills.
  • Ability to work independently and take ownership of technical challenges.
  • High attention to detail and commitment to data accuracy.
Nice to Have:
  • Experience working with government, regulatory, compliance, or public‑sector datasets.
  • Experience with Playwright, Selenium, Puppeteer, Scrapy, or similar scraping frameworks.
  • Experience with data versioning, change detection, or document lineage.
  • Familiarity with AI‑assisted development tools and workflows.
Notes:
  • Years of experience are flexible; problem‑solving ability is more important than seniority.
  • The client values correctness and attention to detail over speed.
  • Candidates should proactively identify and report inconsistencies, missing information, or edge cases.
  • Candidates should be comfortable using AI‑assisted coding tools when appropriate.
  • This is not a customer‑facing role.
  • Technical evaluation may include a live technical discussion, code review, or problem‑solving exercise.
Skills:
  • Web Scraping
  • Data Scraping
  • Data Extraction
  • Python
  • HTML Parsing
  • APIs
  • HTTP Requests
  • FTP
  • Data Transformation
  • Data Normalization
  • Automation
  • ETL
  • PostgreSQL
  • Git
  • Playwright
  • Selenium
  • Puppeteer
  • Scrapy
  • Structured Data
  • Unstructured Data
  • Data Validation
  • Government Data
  • Public Data Sources
  • Technical Documentation
  • Problem Solving
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Data Scraping Engineer: Public Data Pipelines
Data Scraping Engineer: Public Data Pipelines

Trabajosihay • Colombia

Híbrido
Python Developer
Python Developer

Globant • Colombia

A distancia
COP 169.332.000 - 244.591.000
Comprehensive benefits package
Diversity and inclusion commitment
Senior Java/Node Software Engineer
Senior Java/Node Software Engineer

Growth Acceleration Partners • Envigado

Presencial
COP 120.000.000 - 180.000.000
Senior Data Engineer ID71670
Senior Data Engineer ID71670

AgileEngine • Bogotá

Híbrido
COP 281.171.000 - 406.136.000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Data Engineer ID71670
Senior Data Engineer ID71670

AgileEngine • Sur

Híbrido
COP 120.000.000 - 210.000.000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Data Engineer ID71670
Senior Data Engineer ID71670

AgileEngine • Medellín

Híbrido
COP 281.171.000 - 437.377.000
Professional growth
Competitive USD-based pay
Exciting projects
+1
Data Analyst
Data Analyst

Softgic • Medellín

Presencial
COP 103.483.000 - 168.160.000
Data Engineer
Data Engineer

Metova • Colombia

Presencial
COP 30.000 - 45.000
Senior Data Engineer
Senior Data Engineer

Publicis Groupe Holdings B.V • Bogotá

Presencial
COP 219.042.000 - 292.057.000
Data Analyst
Data Analyst

SOFTGIC • Colombia

Presencial
COP 32.000.000 - 52.000.000