SwarmBench Task Engineer (Knowledge/Research) - 75064

Turing

Pune District

On-site

INR 6,704,980 - 9,578,544

Part time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully remote environment
Cutting-edge AI projects
Performance-based contract extension

Job summary

Turing is seeking a highly analytical individual in Pune District, India, to join as a researcher. The role involves crafting challenging problems, analyzing large documents, and designing evaluation benchmarks. Candidates should have a strong research background, proficiency in Python, and experience with JSON. The position offers a fully remote environment and the opportunity to work on cutting-edge AI projects. The engagement is a contractor assignment for 1 month, requiring 40 hours/week with specific PST overlap.

Qualifications

  • 5+ years of research experience in any scientific domain.
  • Ability to extract structured data from unstructured text.
  • Experience with JSON and schema design.
  • Proficiency in Python for data processing.
  • Familiarity with AI coding benchmarks like SWE-bench.
  • Experience with writing Dockerfiles and debugging.

Responsibilities

  • Build benchmark tasks that require analyzing large document collections.
  • Curate research corpora and design comprehensive questions.
  • Write structured ground-truth oracles that can be verified.
  • Design prompts to evaluate agent output against oracles.
  • Create guides for distributing research tasks across agents.

Skills

Research experience
Reading comprehension
Python scripting
Docker
Attention to detail

Tools

JSON
Docker

Job description

Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.

Role Overview

We are seeking a highly analytical and computationally proficient individual to join our team with a strong research background. You will be instrumental in contributing to this role by either crafting challenging and insightful problems in your respective research domain, devising elegant computational solutions.

Responsibilities
  • Build multi-agent benchmark tasks that require reading, analyzing, and synthesizing large document collections
  • Curate real-world research corpora — academic papers, case studies, technical reports — and design questions that require comprehensive analysis
  • Write structured ground-truth oracles (JSON) with specific, verifiable answers that prove the agent actually read the source material
  • Design LLM judge prompts that evaluate agent output field-by-field against the oracle
  • Create decomposition guides that split research across multiple parallel sub-agents (one per document, one per domain, then synthesis)
Required Qualifications
  • 5+ years of research experience (academic or industry) in any scientific domain
  • Strong reading comprehension with ability to extract structured data from unstructured text
  • Experience with JSON and data structures, including schema design and output validation
  • Proficiency in Python scripting for data processing and evaluation (e.g., judge scripts)
  • Familiarity with AI coding benchmarks such as SWE-bench and Terminal-bench
  • Hands-on experience with Docker (writing Dockerfiles, building images, debugging containers)
  • High attention to detail, especially for creating precise evaluation oracles without approximations
Nice to have
  • Experience with systematic reviews, meta-analyses, or large-scale literature surveys
  • Familiarity with medical, legal, or scientific document analysis
  • Experience with NLP or information extraction tasks
  • Knowledge of LLM evaluation and benchmarking (e.g., MMLU, GPQA, SimpleQA)
  • Experience curating datasets for AI evaluation
Perks of Freelancing With Turing
  • Work in a fully remote environment.
  • Opportunity to work on cutting-edge AI projects with leading LLM companies.
  • Potential for contract extension based on performance and project needs.
  • Commitments Required : 40 hours /week with 4 hours of PST Overlap
  • Engagement type : Contractor assignment/freelancer (no medical/paid leave)
  • Duration of contract : 1 month; [expected start date is next week]
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Engineer (Data Analysis)
AI Benchmark Engineer (Data Analysis)

Turing • Delhi

Remote
INR 1,200,000 - 1,800,000
Work on cutting-edge AI projects
Collaborate on high-impact work
Flexible opportunities with global teams
AI Benchmark Software Engineer - 75243
AI Benchmark Software Engineer - 75243

Turing • Hyderabad

Remote
Work on cutting-edge AI projects
Flexible working opportunities
Collaboration with leading organizations
Customer Service Operator - 45024
Customer Service Operator - 45024

Turing • Hyderabad

Remote
Remote-first and flexible work environment
Opportunity to work on cutting-edge AI projects
Potential for contract extension
Remote Data Scientist - 60713
Remote Data Scientist - 60713

Turing • Chennai District

Remote
INR 6,529,000 - 10,262,000
Work in a fully remote environment
Opportunity to work on cutting-edge AI projects
Subject Matter Expert (Law) - 47544
Subject Matter Expert (Law) - 47544

Turing • India

Remote
INR 5,273,000 - 7,910,000
Competitive compensation
Flexible working hours
Opportunity to work on cutting-edge AI projects
+1
Remote LLM Trainer (Agent completion tasks)
Remote LLM Trainer (Agent completion tasks)

Turing • Chennai District

Remote
INR 781,000 - 1,228,000
Fully remote environment
Opportunity to work on cutting-edge AI projects
Flexible hours with a commitment of 20 hours per week
Remote Data Scientist - 60713
Remote Data Scientist - 60713

Turing • India

Remote
INR 5,273,000 - 7,910,000
Fully remote environment
Work on cutting-edge AI projects
Flexible working hours
Remote Data Analyst - 60736
Remote Data Analyst - 60736

Turing • Mumbai

Remote
INR 3,296,000 - 7,910,000
Fully remote environment
Work on cutting-edge AI projects
Flexible commitments
+2
Remote Data Analyst - 60736
Remote Data Analyst - 60736

Turing • Gurugram District

Remote
INR 1,102,000 - 2,204,000
Fully remote environment
Opportunity to work on cutting-edge AI projects
Flexible hours
Remote LLM Trainer (Agent completion tasks)
Remote LLM Trainer (Agent completion tasks)

Turing • Bengaluru

Remote
Fully remote work environment
Work on cutting-edge AI projects
Flexible hours with required daily overlap with PST