Senior Research Engineer, Public Data Benchmarks

Firecrawl

San Francisco (CA)

Hybrid

USD 250.000 - 290.000

Vollzeit

Vor 10 Tagen
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive salary
Equity
PTO
Parental leave
Wellness stipend
Learning budget
Team offsites
Sabbatical

Zusammenfassung

Firecrawl is hiring a Research Engineer - Benchmarks to design, run and publish public benchmarks of data providers, measuring accuracy, coverage, freshness, speed and cost.

You will own datasets, ground truth, scoring, and the weekly release, collaborating with marketing for leaderboards and ensuring results feed back into Alexandria. Requires 4+ years in ML, research engineering, or data engineering with evaluation experience.

Qualifikationen

  • You've shipped evals or benchmarks, and you can explain exactly why your results were trustworthy.
  • You've done this somewhere that matters: a frontier lab, a data company like Scale, Surge, micro1 or Mercor, a third-party benchmark org, or a public benchmark project. OSS contributors very welcome.
  • You're strong in Python and API integrations, and comfortable with messy vendor APIs, rate limits, and inconsistent schemas.
  • You know how to build test sets and scoring methods: sampling, labeling, inter-rater agreement, and when to trust an LLM judge and when not to.
  • You use AI heavily and keep upgrading your own workflow. You ship without waiting for instructions.
  • You write findings clearly for engineers, marketers, and the vendors on the other side of the leaderboard.

Aufgaben

  • Design and run head-to-head benchmarks of data providers against verified ground truth. For example:
  • Build and maintain the test datasets and ground truth, and keep them from going stale or leaking.
  • Own the automated harness and the weekly benchmark release. Every run has to be reproducible, versioned, and defensible.
  • Measure what buyers actually care about: accuracy, coverage, freshness, speed, and cost.
  • Work with marketing to ship public leaderboard pages that are useful and hold up under scrutiny.
  • Close the loop into Alexandria so benchmark results change which provider gets called for what.
  • Keep expanding the categories we benchmark, inside Alexandria and beyond it.

Kenntnisse

Python
API integrations
Benchmarking
Data engineering
Experiment design

Jobbeschreibung

Firecrawl is hiring a Research Engineer - Benchmarks to design, run and publish public benchmarks of data providers, measuring accuracy, coverage, freshness, speed and cost.

You will own datasets, ground truth, scoring, and the weekly release, collaborating with marketing for leaderboards and ensuring results feed back into Alexandria. Requires 4+ years in ML, research engineering, or data engineering with evaluation experience.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Benchmark Engineer — Public Data Provider Leaderboard
Benchmark Engineer — Public Data Provider Leaderboard

Showcify, Inc. • USA

Hybrid
USD 250.000 - 290.000
Salary that makes sense: $250,000–$290
Generous equity
Generous PTO: 15 days +
+5
Research Engineer, Benchmarking
Research Engineer, Benchmarking

Refresh • San Francisco (CA)

Vor Ort
USD 120.000 - 190.000
Research Engineer, Benchmarking
Research Engineer, Benchmarking

Refresh AI • San Francisco (CA)

Vor Ort
USD 120.000 - 150.000
Frontier Benchmarking Research Engineer
Frontier Benchmarking Research Engineer

Refresh • San Francisco (CA)

Vor Ort
USD 120.000 - 190.000
Research Engineer - Benchmarks
Research Engineer - Benchmarks

Showcify, Inc. • USA

Hybrid
USD 250.000 - 290.000
Salary that makes sense: $250,000–$290
Generous equity
Generous PTO: 15 days +
+5
Remote AI Benchmark & Datasets Engineer
Remote AI Benchmark & Datasets Engineer

Ignite Next GmbH • Palo Alto (CA), Northern (KY)

Hybrid
USD 120.000 - 190.000
Remote work
Office visits Palo Alto, Paris, Wroclw
Research Scientist — Frontier Data Benchmarks & AI Evaluation
Research Scientist — Frontier Data Benchmarks & AI Evaluation

Snorkel AI • USA

Vor Ort
USD 200.000 - 350.000
Production ML Engineer – Ranking & Search Systems
Production ML Engineer – Ranking & Search Systems

Mendable • San Francisco (CA)

Hybrid
USD 210.000 - 240.000
Salary that makes sense — $210,000–$0?
Competitive equity — details shared
Generous PTO — 15 days; 24+ days with
+5
Research Engineer - Benchmarks
Research Engineer - Benchmarks

Firecrawl • San Francisco (CA)

Hybrid
USD 250.000 - 290.000
Competitive salary
Equity
PTO
+5
Remote ML Engineer: Build Benchmarks & Production Pipelines
Remote ML Engineer: Build Benchmarks & Production Pipelines

raydar • Northern (KY)

Hybrid
USD 170.000 - 270.000
Competitive equity