Research Engineer, Index Intelligence

Exa

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Exa is seeking a research engineer to own the signals behind the ranking and decision processes that determine what gets indexed and retrieved. You will work across crawling, indexing, and retrieval systems to build end-to-end improvements in search quality.

Ideal candidates have hands-on ML experience with classifiers, rankers, and calibration at web scale and can define success in the absence of ground truth, aligning with Exa's mission to redefine search for the world.

Qualifications

  • Define what the right answer means when ground truth is unavailable.
  • Hands-on ML experience with classifiers, rankers, and calibration at web scale.
  • Ability to think in end-to-end impact and measure improvements in search.

Responsibilities

  • Own the signals behind ranking decisions and the decisions themselves.
  • Collaborate across crawling, indexing, and retrieval teams to integrate your work.
  • Develop models to assess page quality and credibility.
  • Address what a page claims and its suitability for users versus a crawler.

Skills

Ground truth definition
Web-scale ML
Cross-team collaboration
Signal-driven decisions
Page quality modeling

Job description

Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.

Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.

We crawl the web continuously and cannot keep all of it, refresh all of it, or spend a GPU on all of it. Something has to decide which pages are worth indexing, which links are worth following, which pages say the same thing as pages we already have, what has gone stale, and what we should never have picked up. These decisions are worth more than most ranking work. The cheapest way to improve search is to stop indexing pages nobody should ever retrieve, and to start indexing the ones we are missing. Today they run on a mix of learned signals and thresholds someone picked.

We are looking for a research engineer to own the signals behind those decisions, and the decisions themselves.

Desired Experience
  • You are comfortable owning a problem with no ground truth, where the first job is defining what the right answer means

  • Hands-on ML experience with classifiers, rankers and calibration, plus strong data instincts at web scale

  • You think in end-to-end impact. A signal only counts if a decision changes and search gets better

  • You like working across teams. Crawling, indexing and retrieval all consume what you build

  • You care about the problem of finding high quality knowledge and recognize how important this is for the world

Example Projects
  • Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it

  • Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it

  • Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler

  • Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately

Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Index Intelligence
Research Engineer, Index Intelligence

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Engineer, Content Understanding
Research Engineer, Content Understanding

Exa • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer, Content Understanding
Research Engineer, Content Understanding

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research, ML
Research, ML

Exa • San Francisco (CA)

On-site
USD 180,000 - 240,000
Premium healthcare benefits
Fertility benefits
16 weeks parental leave
+1
Research Engineer, Generalist
Research Engineer, Generalist

Exa • San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer, Knowledge Systems
Software Engineer, Knowledge Systems

Exa • San Francisco (CA)

On-site
USD 120,000 - 160,000
Web-Scale ML Research Engineer (Signals & Ranking)
Web-Scale ML Research Engineer (Signals & Ranking)

Exa • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Backend
Software Engineer, Backend

Exa • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Full Stack San Francisco
Software Engineer, Full Stack San Francisco

Exa • San Francisco (CA)

On-site
USD 100,000 - 150,000
Premium healthcare benefits
Fertility benefits
Monthly wellness stipend
Research Engineer, Content Understanding at Web Scale
Research Engineer, Content Understanding at Web Scale

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 240,000