Research Crawling Engineer

Startup Talents

London (KY)

On-site

USD 150,000 - 225,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Startup Talents is hiring a Research Crawling Engineer to join their team in London, United States. This full remote position (USA/EU) entails designing large-scale web data acquisition systems optimized for research and AI model development. Responsibilities include developing fault-tolerant data collection systems, collaborating with research teams, and ensuring data quality. The position offers a competitive salary ranging from $150k to $225k based on experience, along with an equity package. Ideal candidates must possess strong programming skills and experience with web crawlers.

Qualifications

  • Strong programming experience in one or more of Go, Rust, Python, Java, or C++.
  • Experience designing high-throughput, fault-tolerant systems for data collection.
  • Familiarity with NLP pipelines or dataset curation for ML is a plus.

Responsibilities

  • Design high-throughput, fault-tolerant systems for data collection.
  • Collaborate with research teams to align data collection with modeling needs.
  • Monitor crawl performance, coverage, and data quality; iterate quickly.

Skills

Strong programming experience in one or more of: Go, Rust, Python, Java, or C++
Experience building and maintaining large-scale web crawlers or data pipelines
Experience handling anti-bot systems, rate limits, and dynamic/JS-heavy sites
Familiarity with distributed systems and parallel processing
Ability to debug unstable or adversarial environments

Job description

London, United States | Posted on 30/04/2026

They also operate a massive distributed crawler, giving them unique access to high‑quality public web data at global scale.

About the role

They are hiring a Research Crawling Engineer (Full remote - USA/EU, 6 hour overlap with EST). You will join a company at the forefront of developing a web‑scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

As a Research Crawling Engineer, you will design and operate large‑scale web data acquisition systems for research and model development. Your work will span distributed systems, scraping infrastructure, and data pipelines.

Key Responsibilities
  • Operate at the boundary of scale and reliability
  • Adapt to constantly changing web environments
  • Balance throughput, coverage, and data quality
  • Own end‑to‑end data acquisition pipelines
MISSIONS
  • Design high‑throughput, fault‑tolerant systems for data collection (millions to billions of URLs/day)
  • Handle anti‑bot systems, rate limits, and dynamic/JS‑heavy sites
  • Develop pipelines for cleaning, deduplication, filtering, and normalisation
  • Construct and maintain datasets for research and model training
  • Monitor crawl performance, coverage, and data quality; iterate quickly
  • Collaborate with research teams to align data collection with modeling needs
  • Optimize infrastructure for cost, latency, and reliability
Example Projects
  • Build a distributed crawler for a continuously updated, high‑quality web project
  • Design a system to classify and filter billions of pages for pretraining
  • Extract structured data from dynamic, JS‑heavy sites at scale
  • Improve deduplication and quality scoring across multimodal datasets
Requirements
  • Strong programming experience in one or more of: Go, Rust, Python, Java, or C++
  • Experience working for reputable companies
  • Experience building and maintaining large‑scale web crawlers or data pipelines
  • Experience designing high‑throughput, fault‑tolerant systems for data collection (millions to billions of URLs/day)
  • Experience handling anti‑bot systems, rate limits, and dynamic/JS‑heavy sites
  • Experience constructing and maintaining datasets for research and model training
  • Familiarity with distributed systems and parallel processing
  • Experience working with large datasets (TB–PB scale preferred)
  • Ability to debug unstable or adversarial environments
Preferred / Bonus
  • Experience with NLP pipelines or dataset curation for ML
  • Familiarity with LLM pretraining data or retrieval systems
  • Knowledge of proxy systems, IP rotation, and large‑scale request orchestration
  • Background in data quality evaluation or benchmarking
  • Experience running workloads on cloud or bare‑metal infrastructure
Main Evaluation Criteria
  • Ability to design systems that scale without degrading quality
  • Practical problem‑solving under real‑world constraints
  • Speed of iteration and ownership
  • Measurable improvements in data coverage, quality, or efficiency
  • Contract: Permanent role (Full remote – USA or 6 hour overlap with EST)
  • Salary: $150k to $225k based on experience and demonstrated ability to operate at scale + equity package/tokens
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Research Crawling Engineer - Scale Data Pipelines
Remote Research Crawling Engineer - Scale Data Pipelines

Startup Talents • London (KY)

On-site
USD 150,000 - 225,000
Member of Technical Staff - Web Crawl Engineer
Member of Technical Staff - Web Crawl Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Top-tier compensation
Stock options
Health & wellness
+5
Director of Software Engineering (Node.js & Web Scraping Expert)
Director of Software Engineering (Node.js & Web Scraping Expert)

Roman Health Pharmacy LLC • Los Angeles (CA)

On-site
USD 150,000 - 200,000
Health insurance
Flexible working hours
Professional development opportunities
Senior Data Engineer
Senior Data Engineer

Rachel Paul Recruiting • New York (NY)

Hybrid
USD 120,000 - 150,000
Senior Data Engineer - Web Scraping
Senior Data Engineer - Web Scraping

Triwill Group • Spain (TX)

On-site
USD 75,000 - 110,000
Fully remote
Full-time
Autonomy and ownership
+1
Research Engineer — Search/IR
Research Engineer — Search/IR

firecrawl • San Francisco (CA)

On-site
USD 180,000 - 270,000
Competitive salary
Equity options
Generous PTO
+5
Ruby Engineer - Web Scraping (Remote)
Ruby Engineer - Web Scraping (Remote)

SearchApi • Town of Poland (NY)

Remote
USD 100,000 - 130,000
Equity share
Profit sharing
Annual team retreats
Search Engineer
Search Engineer

Firecrawl • United States

Hybrid
USD 190,000 - 260,000
Competitive equity
Generous PTO
Parental leave
+4
Software Engineer III
Software Engineer III

Babel Street • Reston (VA)

Hybrid
USD 110,000 - 135,000
Health benefits covering 85-100% of premiums
Traditional and Roth 401(K) with matching
Unlimited Flexible Leave
+2
Search Engineer
Search Engineer

Firecrawl • San Francisco (CA)

On-site
USD 190,000 - 260,000
Salary that makes sense
Equity up to 0.05%
Generous PTO
+5