Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Three Across is seeking a Subject Matter Expert for Web Scraping and Crawling Web Data Platform to own and extend a production-bound POC. This role combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent in healthcare-data contexts.
Responsibilities include replacing SERPs with paid APIs, async Python pipelines, proxy rotation, and per-domain governance. 8+ years of data acquisition experience and leadership are preferred.
Job Description: Subject Matter Expert Web Scraping and Crawling Web Data Platform
Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Area
Technologies / Capabilities
HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings
[Must]
urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl
[Must]
Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting
[Must]
Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis
[Must]
API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness
[Must]
PDF (pdfplumber, PyMuPDF in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes
[Must]
Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries
[Must]
Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful
[Must]
vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics
[Must]
Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk
[Preferred]
robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal
[Must]
8+ years in data acquisition; 5+ years owning a scraping platform end to end.
Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
Mentored engineers; set crawl standards, review practice, and on-call runbooks.
Degree optional equivalent practical experience is fully accepted.
US payer policy, formulary, or prior-authorization document domain knowledge.
Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
LLM-assisted extraction at controlled cost we run AWS Bedrock in verifier/.
Compliance or legal-review exposure on data acquisition programmes.