Python Data Extraction Engineer

Vserve

Coimbatore District

On-site

INR 900,000 - 1,300,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Vserve in India is seeking a Python Data Extraction Engineer to build and maintain the data-acquisition layer of our platform. You will design robust pipelines to collect and normalize data from government portals, public websites, PDFs, and APIs, handling poorly structured sources at scale.

You will work with GIS, product and engineering teams to validate data quality, detect schema changes, and expose clean datasets for downstream applications.

Qualifications

  • Proficient in Python for data extraction and automation.
  • Experience with web scraping/crawling from diverse sources.
  • Familiar with government portals, PDFs and APIs.
  • Able to handle poorly structured data and changing schemas.
  • Strong QA and data cleaning skills.

Responsibilities

  • Identify and evaluate government/public data sources relevant to property, businesses, and activity.
  • Build Python-based crawlers and extraction pipelines for government portals and public websites.
  • Extract structured information from HTML pages, tables, PDFs, and public APIs.
  • Work with JavaScript-rendered sites and multi-step public search interfaces.
  • Automate recurring extraction with respect to access rules and rate limits.
  • Clean, standardize, and normalize data from different government authorities.
  • Perform entity matching and record linkage using owner, company, address, survey, coordinates, project info.
  • Develop mechanisms to detect website/schema changes and extraction failures.
  • Build validation and QA processes to measure completeness and accuracy.
  • Store extracted information in structured databases and expose clean datasets to downstream apps.
  • Collaborate with GIS, product, and engineering teams to combine location-based signals with records.
  • Research new public/open-data sources to improve accuracy.

Skills

Python
Web scraping
Requests / HTTP clients
BeautifulSoup / lxml

Tools

BeautifulSoup
lxml
Requests

Job description

We are building a data intelligence platform that looks to analyze fragmented public information into structured, actionable business data.

We are looking for a strong Python Data Extraction Engineer who can discover, extract, clean, normalize and integrate data from government portals, public websites, PDFs, APIs and other open data sources.

This is not a conventional application-development role. The ideal candidate enjoys solving difficult data-acquisition problems involving poorly structured websites, inconsistent government portals, changing schemas, PDFs, JavaScript-rendered pages and large volumes of semi-structured information.

What You Will Own

You will build and maintain the data-acquisition layer of the platform.

Key responsibilities include:
  • Identify and evaluate government and public data sources relevant to property, businesses and commercial activity.
  • Build Python-based crawlers and extraction pipelines for government portals and public websites.
  • Extract structured information from HTML pages, tables, PDFs, downloadable files and publicly accessible APIs.
  • Work with JavaScript-rendered websites and multi-step public search interfaces.
  • Automate recurring extraction from multiple sources while respecting applicable access rules, rate limits, and terms.
  • Clean, standardize and normalize inconsistent data from different government authorities.
  • Perform entity matching and record linkage across datasets using fields such as owner name, company name, address, survey number, coordinates and project information.
  • Develop mechanisms to detect website/schema changes and extraction failures.
  • Build validation and QA processes to measure completeness and accuracy.
  • Store extracted information in structured databases and expose clean datasets to downstream applications.
  • Work closely with GIS, product and engineering teams to combine location-based signals with government/public records.
  • Research new public and open-data sources that can improve the accuracy and completeness of our intelligence.
Examples of Data Sources

The work may involve sources such as:

  • Municipal corporation portals
  • State planning and development authorities
  • DTCP and similar planning authorities
  • RERA databases
  • Building and planning permission records
  • Land and property records
  • Tender and procurement portals
  • Company/business registries
  • Government open-data portals
  • Environmental and regulatory approvals
  • Public notices and downloadable government documents
  • Maps and geospatial datasets
  • Other legally accessible public and open-source information
Required Technical Skills

Strong hands-on experience with:

  • Python
  • Web scraping and crawling
  • Requests / HTTP clients
  • BeautifulSoup / lxml
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Python Data Extraction Engineer
Python Data Extraction Engineer

Zohorecruit • Coimbatore District

On-site
INR 900,000 - 1,300,000
Software Engineer, Data Infrastructure and Acquisition
Software Engineer, Data Infrastructure and Acquisition

Analogy Group • India

On-site
INR 3,500,000 - 6,000,000
Senior Data Engineer (Web Scraping)
Senior Data Engineer (Web Scraping)

Oxford Data Plan Ltd. • Chennai District

On-site
INR 2,500,000 - 4,000,000
Senior Data Engineer Web Scraping
Senior Data Engineer Web Scraping

Oxford Data Plan • Indore District

On-site
INR 1,500,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

Foss United • New Delhi

On-site
INR 900,000 - 1,500,000
Data Acquisition Engineer - Python
Data Acquisition Engineer - Python

Eighty Days • Mumbai

On-site
INR 1,200,000 - 1,800,000
Lead Web Scraping Engineer
Lead Web Scraping Engineer

Solytics Infotech • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Python Developer - Web Scraping & Data Crawling
Senior Python Developer - Web Scraping & Data Crawling

Globaldata • Hyderabad

On-site
INR 1,000,000 - 2,000,000
Python Data Engineer (Blr/Chn/Hyd/Kochi/Kol/Pune)
Python Data Engineer (Blr/Chn/Hyd/Kochi/Kol/Pune)

Tata Consultancy Services • Hyderabad, Pune District, Bengaluru

On-site
INR 1,500,000 - 2,600,000
Python ETL Developer
Python ETL Developer

Lonvec Technologies Private Limited • Pune District

On-site
INR 1,000,000 - 1,600,000