Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)

Proof of Skill

Hyderabad

Hybrid

INR 3,500,000 - 5,500,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Chryselys is seeking a senior scraping and data-engineering expert to own and extend a production-bound PoC. You will integrate resilient web-crawling with an AWS Bedrock-based verifier in a healthcare-data-compliance context, while modernizing the stack and ensuring robust observability and cost control.

You will replace SERP scraping with paid APIs, migrate to async workflows, implement per-domain concurrency and proxy strategies, and build tests, alerts, and a licensing register to meet

Qualifications

  • 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
  • Proven scale: 1,000+ domains or 10M+ pages per month in production.
  • Experience rescuing brittle legacy scrapers with measurable cost/reliability gains.
  • Mentored engineers; defined crawl standards, reviews, and on-call runbooks.
  • Equivalent practical experience accepted; degree optional.

Responsibilities

  • Replace scraped SERPs with paid search API behind an existing interface.
  • Migrate fetch.py to async httpx with per-domain concurrency limits.
  • Implement proxy rotation, egress management and multi-IP strategies.
  • Improve failure signaling with typed outcomes, logs, metrics, and alerts.
  • Reverse-engineer payer endpoints and reduce Playwright reliance.
  • Respect robots.txt and per-domain terms; build a licensing register.
  • Containerize and schedule the pipeline; enable CI for offline tests.
  • Add HAR replay and golden-file tests on real payer HTML; run daily canaries.

Skills

Legacy stack comprehension
Reverse engineering
Testing & observability
Mentoring & leadership
Legal & ethical data acquisition

Tools

BeautifulSoup
lxml
XPath/XSLT
Scrapy
Selenium
Playwright
Pydantic validation
Postgres
MongoDB
S3
Airflow/Prefect/Dagster
Docker/K8s
mitmproxy/Charles
OCR/LLM-assisted extraction
pdfplumber/PyMuPDF

Job description

Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.

Role Summary

Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.

Responsibilities
  • Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
  • Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
  • Introduce proxy rotation and egress management; retire the single-IP failure mode.
  • Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
  • Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
  • Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
  • Containerize and schedule the pipeline; add CI running the offline tests on every change.
  • Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Skills

Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]

Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]

Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]

Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]

Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]

Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]

Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]

Experience — Required
  • 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
  • Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
  • Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
  • Mentored engineers; set crawl standards, review practice, and on-call runbooks.
  • Degree optional — equivalent practical experience is fully accepted.
Nice-to-Have
  • US payer policy, formulary, or prior-authorization document domain knowledge.
  • Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
  • LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/.
  • Compliance or legal-review exposure on data acquisition programmes.

Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Consultant - Web Scraping
Senior Consultant - Web Scraping

Keka Technologies Private Limited • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Fullstack Engineer
Fullstack Engineer

Scrapingdog • Rajasthan

On-site
INR 1,800,000 - 2,800,000
Senior Associate - Market Access
Senior Associate - Market Access

Proof of Skill • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Web Crawling Engineer
Senior Web Crawling Engineer

Hexanix One • Ahmedabad District

Hybrid
INR 1,500,000 - 2,100,000
Hybrid work model with flexible hours
Professional development opportunities
Generous PTO and holidays
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)

FullStackTechies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Web Scraping Engineer in Pune
Web Scraping Engineer in Pune

Krawlnet Technologies Pvt Ltd. • Maharashtra

On-site
INR 900,000 - 1,400,000
Senior Web Scraping Engineer
Senior Web Scraping Engineer

Alternative Path • India

On-site
INR 1,200,000 - 1,800,000
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)
Full-Stack Web Scraping Engineer (Python & JavaScript – 3 to 8 Years)

AIMLEAP Inc • India

Remote
INR 1,000,000 - 1,500,000
Consultant - Gen AI Architect (Senior RAG - Document AI Engineer)
Consultant - Gen AI Architect (Senior RAG - Document AI Engineer)

Keka Technologies Private Limited • Hyderabad

On-site
INR 600,000 - 1,200,000
Consultant - Semantic Engineer (Graph RAG - Knowledge Engineer)
Consultant - Semantic Engineer (Graph RAG - Knowledge Engineer)

Keka Technologies Private Limited • Hyderabad

On-site
INR 3,500,000 - 7,000,000