SME-Web Scraping and Crawling — Web Data Platform

Three Across

Hyderabad

On-site

INR 2,500,000 - 4,500,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Three Across is seeking a Subject Matter Expert for Web Scraping and Crawling Web Data Platform to own and extend a production-bound POC. This role combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent in healthcare-data contexts.

Responsibilities include replacing SERPs with paid APIs, async Python pipelines, proxy rotation, and per-domain governance. 8+ years of data acquisition experience and leadership are preferred.

Responsibilities

  • Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
  • Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
  • Introduce proxy rotation and egress management; retire the single-IP failure mode.
  • Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
  • Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
  • Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
  • Containerize and schedule the pipeline; add CI running the offline tests on every change.
  • Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.

Job description

Job Description: Subject Matter Expert Web Scraping and Crawling Web Data Platform

Role summary:

Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities
  • Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
  • Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
  • Introduce proxy rotation and egress management; retire the single-IP failure mode.
  • Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
  • Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
  • Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
  • Containerize and schedule the pipeline; add CI running the offline tests on every change.
  • Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Skills

Area

Technologies / Capabilities

Web & protocol fundamentals

HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings

[Must]

Legacy stack (real mileage)

urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl

[Must]

Modern stack

Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting

[Must]

Reverse engineering

Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis

[Must]

Methodology breadth

API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness

[Must]

Non-HTML extraction

PDF (pdfplumber, PyMuPDF in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes

[Must]

Anti-bot & reliability

Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries

[Must]

Data engineering

Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful

[Must]

Testing & observability

vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics

[Must]

Build vs. buy

Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk

[Preferred]

Legal & ethical

robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal

[Must]

Experience - Required

8+ years in data acquisition; 5+ years owning a scraping platform end to end.

Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.

Rescued a brittle legacy scraper, with before/after reliability and cost numbers.

Mentored engineers; set crawl standards, review practice, and on-call runbooks.

Degree optional equivalent practical experience is fully accepted.

Nice-to-Have

US payer policy, formulary, or prior-authorization document domain knowledge.

Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.

LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.

Compliance or legal-review exposure on data acquisition programmes.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)
Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)

Proof of Skill • Hyderabad

On-site
INR 3,500,000 - 5,500,000
Senior Consultant - Web Scraping
Senior Consultant - Web Scraping

Keka Technologies Private Limited • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer (Web Scraping)
Senior Data Engineer (Web Scraping)

Oxford Data Plan Ltd. • Chennai District

On-site
INR 2,500,000 - 4,000,000
Senior Web Crawling Engineer
Senior Web Crawling Engineer

HexanixOne • Ahmedabad District

On-site
INR 1,500,000 - 2,100,000
Hybrid work model with flexible hours
Professional development opportunities
Generous PTO and holidays
Software Engineer, Data Infrastructure and Acquisition
Software Engineer, Data Infrastructure and Acquisition

Analogy Group • India

On-site
INR 3,500,000 - 6,000,000
Python Mid/Senior Developer – Web Scraping & Automation
Python Mid/Senior Developer – Web Scraping & Automation

Actowiz Solutions LLP • India

On-site
INR 1,200,000 - 1,800,000
Python Mid Senior Developer Web Scraping And Automation
Python Mid Senior Developer Web Scraping And Automation

Hr Actowiz Solutions • Ahmedabad District

On-site
INR 900,000 - 1,500,000
Python Mid/Senior Developer – Web Scraping & Automation
Python Mid/Senior Developer – Web Scraping & Automation

Hr Actowiz Solutions • Ahmedabad District

On-site
INR 1,400,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

Nextyn • Mumbai

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Challenging projects
High growth potential
Python Developer
Python Developer

Resources Valley • Jaipur

On-site
INR 1,200,000 - 1,800,000