Member of Technical Staff (Answer Quality & Evals)

Perplexity

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Perplexity is seeking an engineer for the Answer Quality team in San Francisco to strengthen evaluation foundations for prompts, tools, and models. You will help design and operate the infrastructure that measures agent performance and guides product improvements.

You will collaborate with data scientists, engineers, and product partners to turn evaluation findings into reliable, scalable solutions and to push forward production-grade systems.

Qualifications

  • 4+ years of software, data, or ML engineering shipping and operating production systems.
  • Strong proficiency in Python and SQL, with fundamentals in system design, data modeling, and distributed systems.
  • Experience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelines.
  • Demonstrated ownership of ambiguous technical projects from initial design through production operation.
  • Ability to work effectively with data scientists, engineers, and product partners.

Responsibilities

  • Build shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisions
  • Develop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failures
  • Build and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation data
  • Partner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvements
  • Operate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer Quality

Skills

Python
SQL
System design
Data modeling
Distributed systems
Production systems
Collaboration
Ownership

Tools

Databricks
Snowflake
ClickHouse

Job description

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. The Answer Quality team ensures that our prompts, tools, search systems, datasets, and models work together to create the best possible experience for our users.

As our product and agent capabilities evolve, we need evaluation systems that are fast, reliable, production-faithful, and actionable. In this role, you will build and improve the technical foundations that support Answer Quality across Perplexity. This includes our shared evaluation infrastructure and the platform used to replay and analyze agent traces. You will work closely with data scientists, engineers, and product teams to identify quality problems, measure their impact, and turn evaluation findings into product improvements.

Responsibilities
  • Build shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisions

  • Develop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failures

  • Build and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation data

  • Partner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvements

  • Operate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer Quality

Qualifications
  • 4+ years of software, data, or machine learning engineering experience shipping and operating production systems

  • Strong proficiency in Python and SQL, with solid fundamentals in system design, data modeling, and distributed systems

  • Experience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelines

  • Demonstrated ownership of ambiguous technical projects from initial design through production operation

  • Ability to work effectively with data scientists, engineers, and product partners

Preferred Qualifications
  • Experience building evaluation, experimentation, observability, or machine learning infrastructure

  • Familiarity with LLM and agent systems, including tool use, execution traces, replay, and simulation

  • Experience building on top of large-scale data processing platforms such as Databricks, Snowflake, or ClickHouse

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Answer Quality & Evals)
Member of Technical Staff (Answer Quality & Evals)

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff (Answer Quality & Evals)
Member of Technical Staff (Answer Quality & Evals)

Perplexity • Palo Alto (CA)

On-site
USD 200,000 - 350,000
Member of Technical Staff (Data Scientist, Evals)
Member of Technical Staff (Data Scientist, Evals)

Perplexity • United States

Hybrid
USD 100,000 - 150,000
Staff Engineer, Evaluation & LLM Infrastructure
Staff Engineer, Evaluation & LLM Infrastructure

Perplexity • Palo Alto (CA)

On-site
USD 200,000 - 350,000
Staff Engineer, AI Evaluation & Observability
Staff Engineer, AI Evaluation & Observability

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Staff Engineer - Answer Quality & Evaluation Platforms
Staff Engineer - Answer Quality & Evaluation Platforms

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Data Scientist, AI Evaluation & LLM Quality
Staff Data Scientist, AI Evaluation & LLM Quality

Perplexity • United States

Hybrid
USD 100,000 - 150,000
Member of Technical Staff (Software Engineer, Data Platform)
Member of Technical Staff (Software Engineer, Data Platform)

Perplexity • San Francisco (CA)

On-site
USD 130,000 - 170,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Perplexity • New York (NY), Northern (KY)

On-site
USD 150,000 - 190,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000