Research Scientist – Frontier Evaluations

AfterQuery

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health Insurance
401(k) with Employer Match
Daily Meals: UberEats stipend
Monthly Wellness Stipend

Job summary

AfterQuery is hiring Research Scientists to design and publish rigorous evaluations for frontier AI systems. The role spans agentic, coding, and safety evaluations, across healthcare, STEM, finance, and related fields.

You will own evaluation development end to end and collaborate across disciplines to turn important capability gaps into rigorous public research. PhD in a related field is preferred, and candidates should have a track record of publishing research.

Qualifications

  • Proven track record of publishing benchmarks or research papers.
  • Excellent scientific writing and clear technical communication.
  • Commitment to experimental rigor including baselines, ablations, and statistical validity.
  • Ability to scope ambiguous evaluation questions into reproducible public releases.
  • Experience with agentic, coding, and safety evaluations or applied ML in expert domains.

Responsibilities

  • Lead end-to-end design and validation of frontier AI benchmarks.
  • Collaborate with researchers and domain experts to cover high-priority domains.
  • Analyze model capabilities and failure modes using rigorous experimental methods.
  • Build reproducible evaluation systems, harnesses, graders, and infrastructure.
  • Post-train models and measure resulting performance gains; communicate results via reports and papers.

Skills

Benchmark research
Technical writing
Experimental design
Ambiguity handling
Agentic/safety evals

Education

PhD preferred

Job description

About AfterQuery

AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models. In doing so, we can make expertise that once took a lifetime to build available to anyone who needs it.

Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI. Forbes reported that we could be YC's fastest unicorn, reportedly raising at a $3.2 billion valuation. We're based in San Francisco and backed by leading investors including Altos Ventures, BoxGroup, and Y Combinator and angels from Google DeepMind, OpenAI, Anthropic, Meta Superintelligence Labs, and Microsoft AI.

Why Apply

Massive Opportunity: Forbes reported that we could be YC's fastest unicorn, reportedly raising at a $3.2 billion valuation, and we're not slowing down.

Founding Impact: You will own and architect core infrastructure systems that power our platform from the ground up.

Equity & Growth: Competitive salary and meaningful equity. As we scale, you’ll have the opportunity to shape the engineering organization and lead major technical initiatives.

Strong Team: Our founding team has experience from Citadel Securities, Meta, Google, Silver Lake, and Morgan Stanley — work alongside world-class engineers and researchers.

Overview

AfterQuery is hiring Research Scientists to design and publish rigorous evaluations for frontier AI systems. The role spans agentic, coding, and safety evaluations, as well as expert-domain evaluations involving applied AI in healthcare, STEM, finance, and related fields. You will own evaluation development end to end and collaborate across disciplines to turn important capability gaps into rigorous public research.

Responsibilities
  • Lead the end-to-end design, validation, launch, and continuous improvement of frontier AI benchmarks.

  • Partner with researchers and domain experts to develop evaluations around meaningful model failures, gaps in existing coverage, and high-priority domains.

  • Analyze model capabilities and failure modes using rigorous experimental design and statistical methods.

  • Build reproducible evaluation systems, including harnesses, graders, and benchmark infrastructure.

  • Collaborate with researchers to post-train models and measure the resulting performance gains.

  • Communicate results through benchmark reports, technical articles, and research papers.

Required Qualifications
  • Strong record of publishing benchmarks or research papers.

  • Clear technical communication and strong scientific writing skills.

  • Commitment to experimental rigor, including baselines, ablations, statistical validity, and contamination controls.

  • Ability to take an ambiguous evaluation question from initial scoping through a reproducible public release.

  • Depth in agentic, coding, and safety evaluations or applied machine learning in an expert domain.

Preferred Qualifications
  • PhD in a related technical field.

  • Research publications at leading conferences or peer-reviewed journals.

  • Interest in multidisciplinary research and the creativity to combine methods and insights from AI, engineering, science, and other expert domains.

Company Benefits (For Eligible Employees):
  • Health Insurance: Medical, Vision, Dental

  • 401(k) with Employer Match

  • Daily Meals: Daily UberEats Stipend

  • Monthly Wellness Stipend

We are an equal opportunity employer committed to providing a workplace free from discrimination and harassment. Employment decisions are made without regard to legally protected characteristics under applicable federal, state, or local law.

We comply with applicable pay transparency requirements and provide compensation ranges based on the position, qualifications, experience, and other relevant factors. Reasonable accommodations are available to qualified individuals with disabilities and for sincerely held religious beliefs, as required by law. This job description is intended to describe the general nature and level of work performed and is not an exhaustive list of all duties, responsibilities, qualifications, or working conditions associated with the position. We reserve the right to modify this job description as business needs change.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scholar
Research Scholar

AfterQuery • San Francisco (CA)

On-site
USD 90,000 - 130,000
Health Insurance
401(k) with Employer Match
Daily Meals: UberEats stipend
+1
Software Engineer - Platform/Applied AI
Software Engineer - Platform/Applied AI

AfterQuery • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health Insurance
401(k) with Employer Match
Daily UberEats Stipend
+1
SWE (Latin America)
SWE (Latin America)

AfterQuery • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health Insurance
401(k) with Employer Match
Daily Meals: Daily UberEats Stipend
+1
Strategic Projects Lead - Consulting
Strategic Projects Lead - Consulting

AfterQuery • San Francisco (CA)

On-site
USD 120,000 - 180,000
Health Insurance
401(k) Match
Daily meals stipend
+1
AI/ML Research Intern
AI/ML Research Intern

AfterQuery • San Francisco (CA)

On-site
USD 34,000 - 83,000
Health Insurance
401(k) match
Daily UberEats stipend
+2
Engagement Manager
Engagement Manager

AfterQuery • San Francisco (CA)

On-site
USD 120,000 - 190,000
Health Insurance: Medical, Vision, D,
401(k) with Employer Match
Daily Meals: Daily UberEats Stipend
+1
Senior Software Engineer - Infrastructure & Platform
Senior Software Engineer - Infrastructure & Platform

SupportFinity™ • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary
Equity opportunities
Collaboration with top engineers and researchers
Strategic Projects Lead - Healthcare
Strategic Projects Lead - Healthcare

AfterQuery • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health Insurance
401(k) with Employer Match
Daily Meals: UberEats stipend
+1
Research Technical Program Manager
Research Technical Program Manager

AfterQuery • San Francisco (CA)

On-site
USD 120,000 - 180,000
Health Insurance
401(k) with Employer Match
Daily UberEats Stipend
+1
Technical Strategic Projects Lead
Technical Strategic Projects Lead

AfterQuery • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity opportunities
Competitive salary
Work with a renowned team