Web Scraping Specialist

Wynd Labs

United States

Remote

USD 110,000 - 170,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Opportunity
Culture
Compensation

Job summary

Wynd Labs is seeking a Web Scraping Specialist to lead data extraction efforts, optimize scraping pipelines, and analyze data for high‑quality public web datasets.

You will work on a distributed crawler platform with a small team, leveraging Python or JavaScript and libraries like BeautifulSoup, Scrapy, or Selenium, plus asynchronous techniques.

This fully remote role offers a competitive salary, equity, and the chance to shape the data infrastructure powering frontier AI models.

Qualifications

  • Proven ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects.
  • Proficiency in Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.
  • Knowledge of asynchronous programming, multithreading, and distributed scraping.
  • In-depth knowledge of HTML, CSS, JavaScript, and the DOM.
  • Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity.
  • Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.
  • Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale.
  • Active participation in open-source projects related to web scraping, data processing, or similar fields.

Responsibilities

  • Write, test, and refine code that extracts data from various online sources, ensuring reliability and efficiency.
  • Perform data retrieval tasks, handling complexities such as pagination and dynamic content loaded with AJAX.
  • Clean and format extracted data, ensuring it meets quality standards for further analysis or processing.
  • Database management: Store and manage the scraped data in appropriate databases, optimizing for access speed and data integrity.
  • Regularly monitor the scraping processes, identify and resolve any issues to maintain continuous data flow.

Skills

Data extraction
Python/JavaScript
HTML/CSS/DOM
NoSQL databases
ML for data cleaning
Open-source contributions

Tools

BeautifulSoup
Scrapy
Selenium

Job description

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We’re the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role.

We are seeking a Web Scraping Specialist who is proficient and brings significant experience in data extraction and web scraping techniques. You will join a small, specialized team and lead efforts to gather and analyze data, optimize scraping processes, and support our vision for a future where Grass plays a crucial role in transforming internet data accessibility.

Who You Are.
  • Demonstrated ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects.

  • Proficiency in languages such as Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.

  • Knowledge of asynchronous programming, multithreading, and distributed scraping.

  • In-depth knowledge of HTML, CSS, JavaScript, and the Document Object Model (DOM).

  • Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity.

  • Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.

  • Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale.

  • Active participation in open-source projects related to web scraping, data processing, or similar fields.

What You’ll Be Doing.
  • Write, test, and refine code that extracts data from various online sources, ensuring reliability and efficiency.

  • Perform data retrieval tasks, handling complexities such as pagination and dynamic content loaded with AJAX.

  • Clean and format extracted data, ensuring it meets quality standards for further analysis or processing.

  • Database management: Store and manage the scraped data in appropriate databases, optimizing for access speed and data integrity.

  • Regularly monitor the scraping processes, identify and resolve any issues to maintain continuous data flow.

Why Work With Us:
  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We’re a lean team with a high bar. We come to work not to be comfortable, but to find out what we’re capable of and to do work that matters. We’re not calling for people who keep things moving. We’re calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Web Scraper
Web Scraper

Executive Staff Recruiters / ESR Healthcare • United States

Remote
USD 28,000 - 69,000
Remote Web Scraping Lead
Remote Web Scraping Lead

Wynd Labs • United States

On-site
USD 110,000 - 170,000
Equity package
Fully remote
Remote Web Scraping Engineer | Scalable Data Pipelines
Remote Web Scraping Engineer | Scalable Data Pipelines

Wynd Labs • United States

Remote
USD 110,000 - 170,000
Opportunity
Culture
Compensation
Research Crawling Engineer
Research Crawling Engineer

Startup Talents • London (KY)

On-site
USD 150,000 - 225,000
Senior Software Engineer, Data Infrastructure
Senior Software Engineer, Data Infrastructure

United States Digital Space LLC • United States

Remote
USD 140,000 - 220,000
Senior Data Engineer
Senior Data Engineer

Rachel Paul Recruiting • New York (NY)

On-site
USD 120,000 - 150,000
Research Engineer - Web Crawlers
Research Engineer - Web Crawlers

ElevenLabs • Maine

On-site
USD 140,000 - 210,000
Discretionary stipend
Travel stipend
Co-working stipend
+1
Senior Web Crawling & Data Ingestion Engineer
Senior Web Crawling & Data Ingestion Engineer

XYZ Venture Capital • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Senior Data & Python Software Engineer
Senior Data & Python Software Engineer

Ceartas • United States

Remote
USD 120,000 - 160,000
Web Scraping Data Engineer
Web Scraping Data Engineer

Verition Fund Management LLC • New York (NY)

On-site
USD 125,000 - 200,000