Web-Scale Data Systems Engineer

Visa Hunt

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier compensation
Stock options
Health & wellness
Meals
Life & family
Vacation days
Sponsorship support
Team building

Job summary

Reflection is a research lab building open models and tools for everyone. We are seeking an engineer to own and operate web-scale crawling systems that discover, acquire, and process content from across the internet.

You will manage URL discovery, scheduling, distributed crawling, and data delivery at scale, collaborating with researchers to optimize coverage and data quality. You will work with world-class researchers to determine which parts of the web matter most for model performance and

Qualifications

  • Experience building large‑scale web crawling, search indexing, or internet‑scale data collection systems.
  • Strong understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination.
  • Experience with Ray, Spark, Beam, Flink or similar frameworks for large data processing.
  • Familiarity with content extraction, HTML parsing, browser automation, rendering systems and modern web technologies.
  • Experience operating systems that process petabyte-scale datasets.
  • Strong systems engineering skills including reliability, observability, performance optimization and debugging.
  • Experience designing experiments and using data to improve crawl quality, coverage, and efficiency.
  • Excellent communication and ability to reason about system tradeoffs and constraints.

Responsibilities

  • Build and operate web-scale crawling infrastructure capable of collecting data across billions of URLs.
  • Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration systems.
  • Develop distributed crawlers that acquire content while respecting site constraints and operational needs.
  • Build systems for content extraction, rendering, parsing, and normalization across diverse web formats.
  • Improve crawl coverage, freshness, efficiency, and quality through measurement and experimentation.
  • Design infrastructure for large-scale recrawling, change detection, and incremental updates.
  • Develop specialized crawlers for high-value domains and dynamic websites.
  • Analyze crawl performance and web coverage to identify gaps and opportunities.
  • Build observability, monitoring, and reliability systems for large-scale crawl operations.
  • Debug production issues and continuously improve performance, scalability, and resilience.

Skills

Web crawling
Distributed systems
URL frontier management
Scheduling
Content extraction
HTML parsing
Rendering systems
Browser automation
Petabyte-scale datasets

Education

Bachelor's degree in CS or related

Tools

Ray
Spark
Beam
Flink

Job description

Reflection is a research lab building open models and tools for everyone. We are seeking an engineer to own and operate web-scale crawling systems that discover, acquire, and process content from across the internet.

You will manage URL discovery, scheduling, distributed crawling, and data delivery at scale, collaborating with researchers to optimize coverage and data quality. You will work with world-class researchers to determine which parts of the web matter most for model performance and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Web Crawl Engineer
Member of Technical Staff - Web Crawl Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Top-tier compensation
Stock options
Health & wellness
+5
Data Ingestion Engineering Lead: Web Crawling & Pipelines
Data Ingestion Engineering Lead: Web Crawling & Pipelines

Reflection • New York (NY)

On-site
USD 180,000 - 240,000
Top-tier compensation
Stock options
Health & wellness
+5
Research Crawling Engineer
Research Crawling Engineer

Startup Talents • London (KY)

On-site
USD 150,000 - 225,000
Data Ingestion Engineer for Scalable AI Pipelines
Data Ingestion Engineer for Scalable AI Pipelines

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Stock options
Health & wellness
Meals provided
+1
Remote Research Crawling Engineer - Scale Data Pipelines
Remote Research Crawling Engineer - Scale Data Pipelines

Startup Talents • London (KY)

On-site
USD 150,000 - 225,000
Senior Search Infrastructure Engineer — Scale & Latency
Senior Search Infrastructure Engineer — Scale & Latency

Mendable • San Francisco (CA)

Hybrid
USD 190,000 - 260,000
Equity
Generous PTO
Parental leave
+6
Member of Technical Staff - Data Ingestion Engineer
Member of Technical Staff - Data Ingestion Engineer

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Stock options
Health & wellness
Meals provided
+1
Senior Search Engineer: Build Web Crawling & Ranking
Senior Search Engineer: Build Web Crawling & Ranking

Firecrawl • United States

Hybrid
USD 190,000 - 260,000
Competitive equity
Generous PTO
Parental leave
+4
Remote Senior Data Infrastructure Engineer: Web Crawling
Remote Senior Data Infrastructure Engineer: Web Crawling

ZoomInfo • United States

On-site
USD 140,000 - 220,000
Senior Software Engineer - Scale Web Data Platform in NYC
Senior Software Engineer - Scale Web Data Platform in NYC

String • New York (NY)

On-site
USD 140,000 - 210,000