Senior Data Engineer - Web Scraping

Jobgether

México

Presencial

MXN 800.000 - 1.400.000

Jornada completa

Hace 9 días

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Fully remote position
Full-time opportunity
Autonomy and ownership
Work on advanced web-scraping & data-­
Collaborative engineering team

Descripción de la vacante

Jobgether is assisting a partner to hire a Senior Data Engineer - Web Scraping based in Mexico. This fully remote role focuses on designing and maintaining advanced web scrapers, transforming diverse sources into reliable datasets, and building scalable data pipelines using Python, Pandas, SQL, and Airflow.

The ideal candidate will have 4–6 years of data engineering experience, strong web-scraping expertise (Selenium, Scrapy, XPath), and a solid understanding of HTML/JS/APIs.

Formación

  • Bachelor's or master's degree in Computer Science, Engineering, or a related technical discipline.
  • 4–6 years of professional experience in data engineering or a closely related field.
  • Strong programming skills in Python and strong knowledge of SQL and database technologies.
  • Advanced hands-on expertise with the Python Pandas library for data cleaning, manipulation, exploration, and transformation.
  • Strong web-scraping experience with Selenium, Scrapy, Fiddler, Postman, and XPath.
  • Strong experience with Apache Airflow for workflow orchestration and data pipeline management.
  • Solid understanding of web technologies, including HTML, JavaScript, APIs, and related concepts.
  • Proven experience working with large datasets and performing data cleaning, transformation, manipulation, and replacement.

Responsabilidades

  • Design, develop, deploy, and maintain web scrapers using a range of scraping techniques and tools to collect alternative datasets from diverse sources.
  • Use Python and Pandas to clean, explore, transform, manipulate, and prepare large datasets for downstream consumption.
  • Build and maintain efficient data pipelines that ingest scraped data into databases and data warehouses.
  • Develop and manage scheduled workflows using Apache Airflow and other orchestration tools to ensure reliable and timely data delivery.
  • Collaborate with analysts and cross-functional stakeholders to understand current and anticipated data requirements and translate them into effective technical solutions.
  • Develop quality-control checks to validate data availability, accuracy, consistency, and integrity.
  • Maintain alerting systems, investigate time-sensitive data incidents, and resolve operational issues to ensure reliable day-to-day data delivery.
  • Design and implement tools, applications, and automation that improve the capabilities and efficiency of the web-scraping platform.
  • Contribute to infrastructure and data-product design, bringing practical solutions that support data scientists and other technology teams.
  • Work independently while collaborating with engineering, product, and technology stakeholders to deliver high-quality solutions and continuously improve existing systems.

Conocimientos

Python
SQL
Pandas
Web scraping
Data pipelines
Airflow
Docker
Kubernetes
GitHub Actions
Jenkins
AWS
Selenium
Scrapy
XPath
Postman
Fiddler
APIs
HTML/JavaScript

Educación

Bachelor's or Master's in CS/Engineering

Herramientas

SQL

Descripción del empleo

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer - Web Scraping based in Mexico.

This is a fully remote opportunity for a data engineering professional specializing in web scraping, data processing, and automation.
You’ll design and maintain sophisticated scrapers that transform diverse web-based sources into reliable, high-quality datasets.
Your work will directly support analytical and investment-related decisions by delivering timely data, alerts, and production-ready data products.
The role combines hands-on Python development, data transformation, database engineering, and workflow orchestration.
You’ll collaborate closely with analysts, engineers, and cross-functional teams to understand requirements and build scalable solutions.
With significant ownership and autonomy, you’ll have the opportunity to improve platforms, automate processes, and solve challenging data problems.
The environment is entrepreneurial and team-oriented, with a strong focus on engineering quality, operational reliability, and continuous innovation.

Accountabilities:
  • Design, develop, deploy, and maintain web scrapers using a range of scraping techniques and tools to collect alternative datasets from diverse sources.
  • Use Python and Pandas to clean, explore, transform, manipulate, and prepare large datasets for downstream consumption.
  • Build and maintain efficient data pipelines that ingest scraped data into databases and data warehouses.
  • Develop and manage scheduled workflows using Apache Airflow and other orchestration tools to ensure reliable and timely data delivery.
  • Collaborate with analysts and cross-functional stakeholders to understand current and anticipated data requirements and translate them into effective technical solutions.
  • Develop quality-control checks to validate data availability, accuracy, consistency, and integrity.
  • Maintain alerting systems, investigate time-sensitive data incidents, and resolve operational issues to ensure reliable day-to-day data delivery.
  • Design and implement tools, applications, and automation that improve the capabilities and efficiency of the web-scraping platform.
  • Contribute to infrastructure and data-product design, bringing practical solutions that support data scientists and other technology teams.
  • Work independently while collaborating with engineering, product, and technology stakeholders to deliver high-quality solutions and continuously improve existing systems.
Requirements:
  • Bachelor’s or master’s degree in Computer Science, Engineering, or a related technical discipline.
  • 4–6 years of professional experience in data engineering or a closely related field.
  • Strong programming skills in Python and strong knowledge of SQL and database technologies.
  • Advanced hands-on expertise with the Python Pandas library for data cleaning, manipulation, exploration, and transformation.
  • Strong web-scraping experience with tools and technologies such as Selenium, Scrapy, Fiddler, Postman, and XPath.
  • Strong experience with Apache Airflow for workflow orchestration and data pipeline management.
  • Solid understanding of web technologies, including HTML, JavaScript, APIs, and related concepts.
  • Proven experience working with large datasets and performing data cleaning, transformation, manipulation, and replacement.
  • Ability to design scalable infrastructure, data products, and technical tools for data-focused teams.
  • Strong verbal and written communication skills, with the ability to collaborate effectively with technical and non-technical stakeholders.
  • Self-motivated, detail-oriented, and comfortable working independently while taking ownership of projects and outcomes.
  • Experience with Docker and workload containerization is preferred; Kubernetes experience is a plus.
  • Familiarity with automation and CI/CD technologies such as Jenkins and GitHub Actions is an advantage.
  • Experience with AWS services such as S3, RDS, SNS, SQS, and Lambda is a plus.
Benefits:
  • Fully remote position with the flexibility to work from anywhere.
  • Full-time opportunity within a collaborative, team-oriented engineering environment.
  • Significant autonomy, ownership, and trust in how you approach technical challenges.
  • Opportunity to work on sophisticated web-scraping, data engineering, automation, and data-product initiatives.
  • Exposure to complex datasets supporting analytical and investment-related decision-making.
  • Collaboration with engineering, product, analysts, data scientists, and other technology professionals.
  • Opportunity to contribute to the development and evolution of an entrepreneurial technology team.
  • Professional environment focused on innovation, continuous improvement, and operational excellence.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Data Engineer, Web Scraping (Remote)
Senior Data Engineer, Web Scraping (Remote)

Jobgether • México

A distancia
MXN 800.000 - 1.400.000
Fully remote position
Full-time opportunity
Autonomy and ownership
+2
Web Scraping Engineer
Web Scraping Engineer

QiBit • Ciudad de México

A distancia
Salaire compétitif
Environnement numérique
Culture de travail inclusive
+2
Automation & Integration Engineer
Automation & Integration Engineer

Jobgether • México

A distancia
MXN 6.836.000 - 9.115.000
Remote work in Mexico
Autonomy in role
Junior IT Automation Engineer
Junior IT Automation Engineer

Jobgether • México

A distancia
MXN 558.000 - 1.118.000
Flexible working hours
27 days paid time off
Remote-first culture
+2
Forward Deployed Engineer, AI & Analytics
Forward Deployed Engineer, AI & Analytics

Jobgether • México

A distancia
MXN 900.000 - 1.200.000
Fully remote
Remote engagement with client teams
Opportunities for AI-focused delivery
Data Extractor
Data Extractor

Citian • México

Presencial
Medical, dental, and vision insurance
Generous paid time off
401(k) company match
+2
Senior Data Engineer Id82554
Senior Data Engineer Id82554

Agileengine • Región Centro

Presencial
MXN 900.000 - 1.100.000
Growth opportunities
Competitive compensation
Remote work
+3
Data Engineer (Lead) ID41785
Data Engineer (Lead) ID41785

AgileEngine • Rosarito

Híbrido
MXN 1.531.000 - 2.297.000
Professional growth opportunities
Competitive compensation
Exciting projects
+1
Data Scientist
Data Scientist

S&P Global • Estado de México

Híbrido
MXN 562.000 - 938.000
Be part of a global company
Collaborate with a skilled team
Contribute to high complexity problems
Senior Data Engineer
Senior Data Engineer

Proxify • Ciudad de México

A distancia
MXN 1.469.000 - 2.021.000
Flexible withdrawal options
Up to 24 flex days off per year
Consistent 8-hour working days