Senior Data Engineer (Web Scraping)

Oxford Data Plan Ltd.

Chennai District

On-site

INR 2,500,000 - 4,000,000

Full time

48 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Oxford Data Plan Ltd. is seeking an experienced Data Engineer to own web-scraping and web-data ingestion within our data platform.

You will scale and harden scraping workloads, integrate with our lakehouse architecture, improve scheduling, monitoring and storage, and establish reusable patterns for reliable scrapers. This senior, hands-on role requires independently taking a data source from investigation through deployment and ongoing support, with a track record of production scraping systems

Qualifications

  • Proven experience building production web-scraping systems at scale.
  • Independent ownership of scraping projects from investigation to production.
  • Strong Python engineering for maintainable applications, not just scripts.
  • Experience with scraping tech like Requests/httpx, BeautifulSoup, Scrapy, Playwright or Selenium.
  • Understanding of HTTP/HTML/APIs and browser/network behavior.
  • Experience handling pagination, authentication, rate limits, proxies, and concurrency.
  • Knowledge of data pipelines, data quality, storage and downstream use.

Responsibilities

  • Own web-scraping and data ingestion pipelines.
  • Design scalable scraping framework with patterns for extraction, scheduling, storage, monitoring.
  • Build and maintain production scrapers for new and existing sources.
  • Evaluate build-vs-buy options for scraping infra and services.
  • Investigate sites to choose APIs, HTTP requests, HTML parsing or browser automation.
  • Ensure compliance with policies, robots.txt, privacy and IP constraints.
  • Diagnose issues from changing sites, dynamic content and rate limits.
  • Use AI tooling to improve coding efficiency and maintainability.

Skills

Python engineering
Web scraping at scale
Remote teamwork
Problem solving
Code review & documentation

Tools

Requests/httpx
BeautifulSoup
Scrapy
Playwright/Selenium
Docker
Terraform
Iceberg
PySpark
AWS
Grafana

Job description

We are looking for an experienced, self-driven Data Engineer to take ownership of web scraping and web-based data acquisition within our data platform.

We already use web scraping across a number of ingestion processes, and we are looking to improve the maturity, scale and robustness of this capability. This includes integrating scraping workloads into our lakehouse architecture, improving scheduling, monitoring and storage, and establishing reusable patterns for building and operating scrapers reliably.

This is a senior, hands-on role. We are looking for someone who has previously built and operated production web-scraping systems at scale and can independently take a new data source from investigation through implementation, deployment and ongoing support.

Roles and Responsibilities
  • Own the development and operation of web-scraping and web-data ingestion pipelines.
  • Design and establish a scalable web-scraping framework, including common patterns for extraction, scheduling, storage, monitoring, validation and failure handling.
  • Design, build and maintain reliable production scrapers for new and existing data sources.
  • Evaluate and make build-versus-buy recommendations for scraping infrastructure and third-party services, considering capability, reliability, cost, operational complexity and risk.
  • Investigate websites and determine the most appropriate acquisition approach, including APIs, direct HTTP requests, HTML parsing, browser automation or third-party tooling.
  • Ensure scraping activities appropriately consider internal policies, website terms, robots.txt, access restrictions, privacy and intellectual-property constraints, escalating unclear cases where needed.
  • Diagnose and resolve issues caused by changing websites, dynamic content, authentication, rate limits and other operational challenges.
  • Use AI tooling to work more effectively while understanding, reviewing and being able to defend the code you ship.
Required Qualifications
  • Demonstrated professional experience building and operating production web-scraping systems at scale.
  • Proven ability to independently own substantial scraping projects from initial investigation through production operation.
  • Strong production Python engineering skills, including building maintainable applications rather than standalone scripts.
  • Strong practical experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright or Selenium.
  • Good understanding of HTTP, HTML, APIs, JavaScript-rendered websites and browser/network behaviour.
  • Experience handling common scraping challenges such as pagination, authentication, sessions, retries, rate limiting, concurrency and proxies.
  • Strong understanding of data pipelines, data quality and how collected data should be validated, stored and consumed downstream.
  • Experience deploying, monitoring and supporting production workloads in a cloud environment.
  • Strong debugging and problem-solving skills, with the ability to work independently and make sensible engineering decisions.
  • Experience working effectively within a remote engineering team, including code review, documentation and ticket-based workflows.
Desirable Skills
  • Experience with AWS.
  • Experience with lakehouse or data-lake architectures, particularly Iceberg.
  • Experience with PySpark or other distributed data-processing technologies.
  • Experience with Docker and containerised workloads.
  • Experience with Terraform or other infrastructure-as-code tooling.
  • Experience with Grafana or similar observability platforms.
  • Experience operating high-volume or distributed crawling systems.
  • Experience evaluating or operating commercial scraping, proxy or browser-infrastructure services.
  • Experience implementing automated scraper testing, canary runs or source-drift detection.
  • Experience working with legal, compliance, privacy or data-governance teams on web-data acquisition.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Nextyn • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Challenging projects
High growth potential
SME-Web Scraping and Crawling — Web Data Platform
SME-Web Scraping and Crawling — Web Data Platform

Three Across • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Senior Data Engineer
Senior Data Engineer

Solytics Infotech • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

DigitalXNode • Hyderabad

On-site
INR 1,500,000 - 2,300,000
Senior Web Crawling Engineer
Senior Web Crawling Engineer

HexanixOne • Ahmedabad District

On-site
INR 1,500,000 - 2,100,000
Hybrid work model with flexible hours
Professional development opportunities
Generous PTO and holidays
Data Engineering Manager – Web Crawling & Pipeline Architecture
Data Engineering Manager – Web Crawling & Pipeline Architecture

Outsource Bigdata Solutions • Bengaluru

On-site
INR 1,800,000 - 2,500,000
Flexible working hours
Health benefits
Career development opportunities
Python Mid/Senior Developer – Web Scraping & Automation
Python Mid/Senior Developer – Web Scraping & Automation

Actowiz Solutions LLP • India

On-site
INR 1,200,000 - 1,800,000
Senior Data Engineer
Senior Data Engineer

Toptal • India

On-site
INR 3,000,000 - 6,000,000
Senior Data Engineer (Data & Analytics)
Senior Data Engineer (Data & Analytics)

Summit Consulting Services • Ernakulam

On-site
INR 1,000,000 - 1,800,000
Collaborative culture
Opportunity to influence architecture
Modern cloud-native tech stack
Senior Data Engineer
Senior Data Engineer

Zettamine Labs • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,500,000 - 3,200,000