An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Turing is seeking an investigator to design an evaluation benchmark for frontier AI browsing agents. This role focuses on creating research problems that challenge state-of-the-art systems, starting from verifiable facts and building auditable evidence trails.
You will produce a natural-language research question with a stable answer, clues spanning diverse fact types, and a validation record showing searches and results.
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
We are building an evaluation benchmark for frontier AI browsing agents. Your job is to design research problems that a state-of-the-art AI cannot solve, even with full web access and multiple attempts. This is not a subject matter expert role, nor is it a content-writing role. It is investigative research.
You will start from a verifiable fact, work backwards to construct a question that makes that fact extremely hard to locate, and then prove your work with a complete, auditable evidence trail.
Familiarity with JSON and structured data delivery formats