Research Engineer, Content Understanding

Speedrun Talent Network

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Exa is an applied AI lab building a groundbreaking search engine. We develop massive-scale infrastructure to crawl the web, train state-of-the-art embedding models, and design high-performance vector databases to retrieve information globally.

As a backend engineer, you will contribute to search architecture, work on challenging problems like credibility and misinformation, and help define ground-truth where it does not yet exist.

Qualifications

  • Graduate-level ML degree or equivalent with ≥2 years experience.
  • Ability to build a PyTorch transformer from scratch and scale for cost.
  • Experience building large-scale datasets and supervision-focused research.
  • Comfort with defining ground truth where none exists yet.
  • Interest in high-quality knowledge retrieval and misinformation handling.

Responsibilities

  • Develop search architecture components and experiments.
  • Work on large-scale data and supervision strategies for ML models.
  • Collaborate with researchers and engineers across teams.
  • Prototype and evaluate models for web-scale reliability.

Skills

Graduate-level ML
PyTorch transformer
Large-scale datasets
Ground-truth definition
Knowledge extraction

Education

Master’s or PhD in ML

Job description

Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.

Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.

As a backend engineer, you'd play a critical role in our search architecture. We're pretty flexible on what projects people work on based on their skills and interests.

Search quality is bounded by what we understand about a page. Before anything can be retrieved, something has to work out what the page actually says. That means parsing it into the parts that are content and the parts that are furniture, classifying what kind of page it is and what it is about, telling whether the page is usable at all, extracting when it was published, judging how good it is and whether it can be trusted, and working out whether it says anything that a page we already have does not. All of this has to work on every page on the web, in every language, in every shape the web comes in.

Some of this is classic document understanding. Some of it is much more open. Credibility and misinformation, AI-generated and machine-spun content, and pages written to be found rather than read are all unsolved, and search results are only as trustworthy as our answers to them.

We are looking for a research engineer to work on this. There is a lot of room to do it well.

Desired Experience
  • Graduate-level ML experience (Master’s or PhD with at least 2 years of relevant experience), or an exceptionally strong undergrad

  • You can build a transformer from scratch in PyTorch, and you have trained models that then had to be cheap enough to run everywhere

  • You like building large-scale datasets and living in the data. Most of the wins here are in the supervision rather than the architecture

  • You are comfortable with problems where the ground truth does not exist yet and defining it is part of the job

  • You care about the problem of finding high quality knowledge and recognize how important this is for the world

Example Projects
  • Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it

  • Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it

  • Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler

  • Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately

  • Build classification and extraction that is accurate at web scale and cheap enough to run on all of it

  • Design the supervision for something nobody has labels for, and find out whether it is learnable at all

  • Trace a bad search result back to the page-level prediction that caused it, and fix it at the source

Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Content Understanding
Research Engineer, Content Understanding

Exa • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer, Index Intelligence
Research Engineer, Index Intelligence

Speedrun Talent Network • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Engineer, Index Intelligence
Research Engineer, Index Intelligence

Exa • San Francisco (CA)

On-site
USD 140,000 - 210,000
Research, ML
Research, ML

Exa • San Francisco (CA)

On-site
USD 180,000 - 240,000
Premium healthcare benefits
Fertility benefits
16 weeks parental leave
+1
Software Engineer, Knowledge Systems
Software Engineer, Knowledge Systems

Exa • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Engineer, Content Understanding — Scale Web Knowledge
Research Engineer, Content Understanding — Scale Web Knowledge

Exa • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer, Generalist
Research Engineer, Generalist

Exa • San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer, Backend
Software Engineer, Backend

Exa • San Francisco (CA)

On-site
USD 140,000 - 210,000
Technical Product Marketing, Content Exa AI San Francisco, California
Technical Product Marketing, Content Exa AI San Francisco, California

Neura Market • San Francisco (CA)

On-site
USD 120,000 - 170,000
Health, dental, vision
Parental leave & fertility benefits
Office meals & snacks
+1
Digital Marketing Manager
Digital Marketing Manager

Exa • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health and wellness benefits
Fertility and family planning support
Office amenities and meals