Senior Data Engineer - Web Scraping

Jobgether

India

On-site

INR 1,400,000 - 2,600,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote position
Full-time opportunity
Autonomy and ownership
Challenging projects

Job summary

Jobgether, listed on behalf of a partner company, seeks a Senior Data Engineer - Web Scraping based in India. This fully remote role focuses on designing and maintaining scrapers to turn diverse web sources into reliable datasets, enabling analytical and investment decisions.

You will work with Python, Pandas, SQL, Airflow, and orchestrate pipelines, while collaborating with analysts and engineers. The position offers significant ownership, autonomy, and opportunities to innovate in a

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or related technical discipline.
  • 4–6 years of professional experience in data engineering or closely related field.
  • Strong Python programming and SQL/database knowledge.
  • Advanced Pandas experience for data cleaning and transformation.
  • Strong web-scraping experience with Selenium, Scrapy, Fiddler, Postman, XPath.
  • Experience with Apache Airflow for workflow orchestration.
  • Solid understanding of HTML, JavaScript, APIs and related concepts.
  • Experience with large datasets, data cleaning, transformation, and replacement.
  • Ability to design scalable data infrastructure and tools for data teams.
  • Excellent collaboration and communication skills.

Responsibilities

  • Design, develop, deploy, and maintain web scrapers using diverse techniques.
  • Use Python and Pandas to clean, transform, and prepare large datasets.
  • Build and maintain data pipelines into databases and data warehouses.
  • Develop and manage scheduled workflows with Apache Airflow.
  • Collaborate with analysts and stakeholders to translate requirements into solutions.
  • Implement quality-control checks for data availability and accuracy.
  • Maintain alerting systems and resolve operational data incidents.
  • Create tools and automation to improve web-scraping platform efficiency.
  • Contribute to data-product and infrastructure design for data science teams.
  • Work independently while coordinating with engineering and product teams.

Skills

Python
SQL
Pandas
Web scraping
Airflow
Docker
Kubernetes
AWS
Jenkins
GitHub Actions
REST APIs
HTML/JavaScript

Education

Bachelor's or Master's in CS/Engineering

Tools

Selenium
Scrapy
Fiddler
Postman
XPath
Apache Airflow
Docker
Kubernetes
Jenkins
GitHub Actions
AWS (S3, RDS, SNS, SQS, Lambda)

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer - Web Scraping based in India.

This is a fully remote opportunity for a data engineering professional specializing in web scraping, data processing, and automation.
You'll design and maintain sophisticated scrapers that transform diverse web-based sources into reliable, high-quality datasets.
Your work will directly support analytical and investment-related decisions by delivering timely data, alerts, and production-ready data products.
The role combines hands-on Python development, data transformation, database engineering, and workflow orchestration.
You'll collaborate closely with analysts, engineers, and cross-functional teams to understand requirements and build scalable solutions.
With significant ownership and autonomy, you'll have the opportunity to improve platforms, automate processes, and solve challenging data problems.
The environment is entrepreneurial and team-oriented, with a strong focus on engineering quality, operational reliability, and continuous innovation.

Accountabilities:

  • Design, develop, deploy, and maintain web scrapers using a range of scraping techniques and tools to collect alternative datasets from diverse sources.
  • Use Python and Pandas to clean, explore, transform, manipulate, and prepare large datasets for downstream consumption.
  • Build and maintain efficient data pipelines that ingest scraped data into databases and data warehouses.
  • Develop and manage scheduled workflows using Apache Airflow and other orchestration tools to ensure reliable and timely data delivery.
  • Collaborate with analysts and cross-functional stakeholders to understand current and anticipated data requirements and translate them into effective technical solutions.
  • Develop quality-control checks to validate data availability, accuracy, consistency, and integrity.
  • Maintain alerting systems, investigate time-sensitive data incidents, and resolve operational issues to ensure reliable day-to-day data delivery.
  • Design and implement tools, applications, and automation that improve the capabilities and efficiency of the web-scraping platform.
  • Contribute to infrastructure and data-product design, bringing practical solutions that support data scientists and other technology teams.
  • Work independently while collaborating with engineering, product, and technology stakeholders to deliver high-quality solutions and continuously improve existing systems.
Requirements:
  • Bachelor's or master's degree in Computer Science, Engineering, or a related technical discipline.
  • 4–6 years of professional experience in data engineering or a closely related field.
  • Strong programming skills in Python and strong knowledge of SQL and database technologies.
  • Advanced hands-on expertise with the Python Pandas library for data cleaning, manipulation, exploration, and transformation.
  • Strong web-scraping experience with tools and technologies such as Selenium, Scrapy, Fiddler, Postman, and XPath.
  • Strong experience with Apache Airflow for workflow orchestration and data pipeline management.
  • Solid understanding of web technologies, including HTML, JavaScript, APIs, and related concepts.
  • Proven experience working with large datasets and performing data cleaning, transformation, manipulation, and replacement.
  • Ability to design scalable infrastructure, data products, and technical tools for data-focused teams.
  • Strong verbal and written communication skills, with the ability to collaborate effectively with technical and non-technical stakeholders.
  • Self-motivated, detail-oriented, and comfortable working independently while taking ownership of projects and outcomes.
  • Experience with Docker and workload containerization is preferred; Kubernetes experience is a plus.
  • Familiarity with automation and CI/CD technologies such as Jenkins and GitHub Actions is an advantage.
  • Experience with AWS services such as S3, RDS, SNS, SQS, and Lambda is a plus.
Benefits:
  • Fully remote position with the flexibility to work from anywhere.
  • Full-time opportunity within a collaborative, team-oriented engineering environment.
  • Significant autonomy, ownership, and trust in how you approach technical challenges.
  • Opportunity to work on sophisticated web-scraping, data engineering, automation, and data-product initiatives.
  • Exposure to complex datasets supporting analytical and investment-related decision-making.
  • Collaboration with engineering, product, analysts, data scientists, and other technology professionals.
  • Opportunity to contribute to the development and evolution of an entrepreneurial technology team.
  • Professional environment focused on innovation, continuous improvement, and operational excellence.

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - Web Scraping
Data Engineer - Web Scraping

Alternative Path • India

On-site
INR 900,000 - 1,300,000
Docker-enabled workflows
Senior Web Scraping Engineer
Senior Web Scraping Engineer

Alternative Path • India

On-site
INR 1,200,000 - 1,800,000
Senior Data Engineer
Senior Data Engineer

Nextyn • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Challenging projects
High growth potential
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)

AIMLEAP Inc • India

Remote
INR 1,000,000 - 1,500,000
Senior Web Crawling Engineer
Senior Web Crawling Engineer

Hexanix One • Ahmedabad District

Hybrid
INR 4,000,000 - 6,000,000
Hybrid work model
Professional development
Team building events
+1
Data Engineering Manager – Web Crawling & Pipeline Architecture
Data Engineering Manager – Web Crawling & Pipeline Architecture

AIMLEAP Inc • Bengaluru

Remote
INR 2,000,000 - 3,000,000
Sr. Software Engineer (C++, Python)
Sr. Software Engineer (C++, Python)

Jobgether • India

On-site
INR 2,500,000 - 4,000,000
Stock option offering
401(k) plan with employer matching
Health, dental, vision insurance
Python Developer
Python Developer

Ascendion • Gurugram District

On-site
INR 1,000,000 - 2,000,000
AI Enabled Data Scraping Engineer – Junior
AI Enabled Data Scraping Engineer – Junior

AIMLEAP Inc • Bengaluru

Remote
INR 350,000 - 520,000
Data Engineering Manager – Web Crawling & Pipeline Architecture
Data Engineering Manager – Web Crawling & Pipeline Architecture

Outsource Bigdata Solutions • Bengaluru

Remote
INR 1,800,000 - 2,500,000
Flexible working hours
Health benefits
Career development opportunities