AI-Driven Data Crawling Specialist (Remote)

Rubick

Oregon (WI)

Hybrid

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Work from home - Monday to Saturday

Job summary

Rubick is seeking a data crawling engineer to build and maintain its data acquisition engine for extracting, validating, and structuring product information from eCommerce sites. This role powers Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.

You will develop scalable crawlers using Python and frameworks like BeautifulSoup, Scrapy, Selenium, and Playwright, focusing on AI-assisted extraction, data validation, and high-quality data pipelines.

Qualifications

  • Proficient in Python for web scraping and automation.
  • Hands-on experience with HTML/CSS and DOM structures.
  • Experience with data validation, cleaning, and structuring of datasets.
  • Familiarity with AI-assisted data extraction and ML data pipelines is a plus.
  • Able to handle large datasets while maintaining data quality.
  • Experience in eCommerce, retail, or marketplace data is preferred.

Responsibilities

  • Develop, maintain, and optimize web crawlers to extract product data from eCommerce sites.
  • Collect structured and unstructured product information including details, pricing, images, specs, and availability.
  • Clean, validate, and organize extracted datasets for accuracy and consistency.
  • Monitor crawler performance and resolve extraction issues.
  • Collaborate with Product, Engineering, and Data teams to ensure timely data collection.
  • Maintain documentation for crawling processes, extraction rules, and data quality standards.
  • Leverage AI-powered extraction techniques to improve accuracy and efficiency.
  • Build intelligent crawling workflows using Python and modern web automation frameworks.
  • Implement automated data validation and monitoring systems for high-quality datasets.
  • Collaborate with AI/Data Engineering teams to support ML models and Product Intelligence systems.
  • Continuously optimize crawling performance for speed, scalability, and reliability.
  • Identify opportunities to automate repetitive extraction and validation workflows.
  • Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
  • Build reusable crawling frameworks and standardized extraction pipelines.
  • Contribute to knowledge repositories and process improvements across data acquisition.

Skills

Python for scraping
HTML/CSS/DOM
Data quality & validation
Analytical thinking
Large datasets
E-commerce domain knowledge
AI-assisted extraction exposure

Tools

BeautifulSoup
Scrapy
Selenium
Playwright

Job description

Rubick is seeking a data crawling engineer to build and maintain its data acquisition engine for extracting, validating, and structuring product information from eCommerce sites. This role powers Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.

You will develop scalable crawlers using Python and frameworks like BeautifulSoup, Scrapy, Selenium, and Playwright, focusing on AI-assisted extraction, data validation, and high-quality data pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Crawling - Consultant
Data Crawling - Consultant

Rubick • Oregon (WI)

Hybrid
USD 90,000 - 130,000
Work from home - Monday to Saturday
Research Crawling Engineer
Research Crawling Engineer

Startup Talents • London (KY)

On-site
USD 150,000 - 225,000
Remote Ruby Engineer - Web Scraping & Automation
Remote Ruby Engineer - Web Scraping & Automation

SearchApi • United States

Remote
USD 120,000 - 180,000
Remote Web Scraping Lead
Remote Web Scraping Lead

Wynd Labs • United States

On-site
USD 110,000 - 170,000
Equity package
Fully remote
Remote Ruby Engineer - Web Scraping & Browser Automation
Remote Ruby Engineer - Web Scraping & Browser Automation

SearchApi • Town of Sweden (NY)

On-site
USD 110,000 - 170,000
Fully remote
Equity share
Profit sharing
+1
Remote Senior Data Infrastructure Engineer: Web Crawling
Remote Senior Data Infrastructure Engineer: Web Crawling

ZoomInfo • United States

On-site
USD 140,000 - 220,000
Remote Ruby Engineer: Web Scraping & Automation — Equity
Remote Ruby Engineer: Web Scraping & Automation — Equity

SearchApi • Germany (OH)

On-site
USD 120,000 - 160,000
Fully Remote
Equity share
Profit sharing
+2
Data Ingestion Engineering Lead: Web Crawling & Pipelines
Data Ingestion Engineering Lead: Web Crawling & Pipelines

Reflection • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Automation Engineer - Worlwide Remote
Automation Engineer - Worlwide Remote

Invisible Technologies • San Francisco (CA)

Remote
USD 80,000 - 120,000
Remote Python Data Scraping Engineer for AI Projects
Remote Python Data Scraping Engineer for AI Projects

Mindrift • Michigan

On-site
USD 36,368 - 51,797
Performance-based bonus programs
Fully remote work