Research Engineer, Web Crawling

XYZ Venture Capital

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines in San Francisco seeks a Software Engineer to own our web-crawling systems from distributed collection to data filtering and pretraining data preparation.

You'll build crawlers, ingestion pipelines, and petabyte-scale infrastructure, collaborate with pretraining teams, and help set technical direction while mentoring peers.

Qualifications

  • 8+ years designing, building, and scaling web crawlers or large-scale distributed data-acquisition systems.
  • Own crawler or data-acquisition infrastructure at internet scale.
  • Strong software engineering in Python, Go, or Rust with distributed systems experience.
  • Knowledge of robots.txt, rate limiting, licensing for large-scale data collection.

Responsibilities

  • Design and scale the web crawler and ingestion infrastructure sourcing Inkling's pretraining data.
  • Build pipelines for large-scale extraction, deduplication, and data quality filtering.
  • Build specialized crawlers for high-value data sources.
  • Collaborate with the pretraining team to understand data impact on model quality.
  • Improve reliability and efficiency of crawling at petabyte scale.
  • Set technical direction and mentor engineers.

Skills

Distributed systems
Web crawling
Go
Rust
Python

Tools

Python

Job description

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.

What You’ll Do
  • Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data

  • Build pipelines for large-scale extraction, deduplication, and data quality filtering

  • Build specialized crawlers for high-value or hard-to-reach data sources

  • Work with the pretraining team to understand how changes in crawled data affect model performance

  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale

  • Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned

Skills & Qualifications
Minimum Qualifications
  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems

  • A track record of owning crawler or data-acquisition infrastructure at internet scale

  • Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems

  • Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)

Preferred Qualifications
  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale

  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title

  • Experience designing systems for petabyte-scale storage and processing

  • Track record of open-source contributions to crawling, scraping, or data infrastructure tools

  • Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-$475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Web Crawling & Data Ingestion Engineer
Senior Web Crawling & Data Ingestion Engineer

XYZ Venture Capital • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health benefits
Dental & Vision benefits
Unlimited PTO
+2
Software Engineer, Research Tools
Software Engineer, Research Tools

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Research Software Engineer, Post Training
Research Software Engineer, Post Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 300,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Paid parental leave
+1
Technical Recruiter
Technical Recruiter

Thinking Machines • San Francisco (CA)

On-site
USD 200,000 - 275,000
Health, dental, vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Pre-Training Data
Research, Pre-Training Data

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Mid Training
Research, Mid Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Research Crawling Engineer
Research Crawling Engineer

Startup Talents • London (KY)

On-site
USD 150,000 - 225,000
Software Engineer, Product
Software Engineer, Product

XYZ Venture Capital • San Francisco (CA)

On-site
USD 300,000 - 475,000