Get more replies from employers
Send a job-specific resume in minutes.
Turing is seeking experienced AI Benchmark Engineers to design and develop multi-agent benchmark tasks for evaluating advanced AI systems. The role involves creating benchmark tasks that require AI agents to analyze complex datasets and derive specific conclusions.
Ideal candidates should have 5+ years in data analysis, strong proficiency in SQL and Python, and experience with real-world messy datasets. This is a contractor position for 4 weeks with a commitment of 8 hours per day, allowing for flexible remote work.
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps customers in two ways: working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
We are seeking experienced AI Benchmark Engineers — Data Analysis to design and develop high-quality multi-agent benchmark tasks that evaluate the analytical reasoning, coordination, and execution capabilities of advanced AI systems.
In this role, you will build realistic benchmark tasks that require AI agents to analyze large, messy, multi-source datasets, decompose work across specialist sub-agents, and arrive at specific, verifiable conclusions. These tasks may involve structured and semi-structured data such as CSVs, JSON files, logs, reports, survey results, vendor assessments, or financial and operational documents.
Your work will help measure how effectively AI systems perform complex analytical workflows involving cross-referencing, contradiction detection, anomaly identification, and statistical reasoning across multiple data sources.