AI Benchmark Software Engineer - 75243

Turing

Hyderabad

Remote

INR 6,572,519 - 9,201,526

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Work on cutting-edge AI projects
Flexible working opportunities
Collaboration with leading organizations

Job summary

Turing is seeking experienced Engineers — Code / SWE to design high-quality multi-agent benchmark tasks. You will focus on real open-source code changes like bug fixes and refactors to evaluate AI agent performance.

This remote contractor role requires work on leading AI projects with a commitment of 8 hours per day and entails collaborating with global teams without included medical/paid leave. The contract is expected to last for 4 weeks, with an immediate start.

Qualifications

  • Strong experience navigating frameworks like Django, Flask, FastAPI, Node.js.
  • Knowledge in Docker including writing Dockerfiles.
  • Proven ability to write clear technical specifications.

Responsibilities

  • Build multi-agent benchmark tasks using real-world code changes.
  • Work with the Harbor framework within Docker.
  • Design verification scripts to validate AI-generated code.

Skills

Experience with AI coding benchmarks
Experience reading large open-source codebases
Familiarity with Git workflows
Comfortable working with Docker
Experience writing test scripts
Ability to write technical specifications

Job description

Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps customers in two ways: working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.

Role Overview:

We are looking for experienced Engineers — Code / SWE to design and build high-quality multi-agent benchmark tasks based on real-world software engineering workflows.

In this role, you will create tasks grounded in real open-source code changes such as bug fixes, migrations, and refactors. These tasks are used to evaluate how effectively AI agents can understand large codebases, apply precise modifications, and produce correct, testable outputs.

You will work within a structured evaluation framework (Harbor), define clear task instructions, design verification logic, and decompose complex engineering problems across multiple specialized agents.

What does day-to-day look like:
  • Build multi-agent benchmark tasks based on real-world open-source code changes (bug fixes, migrations, refactors)
  • Work with the Harbor evaluation framework to run and validate tasks inside Docker environments
  • Write clear, precise task instructions specifying file paths, function signatures, expected behavior, and constraints
  • Design and implement Python-based verification scripts to validate correctness of agent-generated code changes
  • Create decomposition strategies that split complex code changes across multiple independent sub-agents
  • Run, debug, and refine tasks within containerized environments to ensure reproducibility and determinism
  • Evaluate task performance signals and improve task quality, clarity, and difficulty
Requirements:
  • Experience with AI coding benchmarks (e.g., SWE-bench, Terminal-Bench)
  • Strong experience reading and navigating large open-source codebases (e.g., Django, Flask, FastAPI, Node.js, or similar)
  • Familiarity with Git workflows, including pull requests, diffs, cherry-picking, and working with specific commits
  • Comfortable working with Docker (writing Dockerfiles, building images, debugging container issues)
  • Experience writing test scripts (pytest, unittest, or custom assertion-based testing)
  • Ability to write clear, precise, and unambiguous technical specifications
  • Perks of Freelancing With Turing
  • Work on cutting-edge AI projects with leading foundation model companies
  • Collaborate on high-impact work at the frontier of LLM evaluation and reasoning
  • Remote, flexible opportunities with global teams
  • Commitments Required: 8 hours per day with a 4-hour overlap with PST.
  • Employment Type: Contractor position (Note: this role does not include medical/paid leave).
  • Duration of Contract: 4 weeks; [expected start date is next week].
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmark Engineer (Data Analysis)
AI Benchmark Engineer (Data Analysis)

Turing • Delhi

Remote
INR 1,200,000 - 1,800,000
Work on cutting-edge AI projects
Collaborate on high-impact work
Flexible opportunities with global teams
SwarmBench Task Engineer (Knowledge/Research) - 75064
SwarmBench Task Engineer (Knowledge/Research) - 75064

Turing • Pune District

On-site
INR 6,704,980 - 9,578,544
Fully remote environment
Cutting-edge AI projects
Performance-based contract extension
Remote Data Engineer - 60713
Remote Data Engineer - 60713

Turing • Pune District

Remote
INR 5,273,000 - 11,864,000
Fully remote environment
Cutting-edge AI projects
Flexible commitment hours
+1
Remote Data Engineer - 60713
Remote Data Engineer - 60713

Turing • Hyderabad

Remote
INR 5,273,000 - 9,228,000
Fully remote work environment
Opportunity to work on cutting-edge AI projects
Flexible hours with at least 20 hours per week
Remote Data Scientist - 60713
Remote Data Scientist - 60713

Turing • Chennai District

Remote
INR 6,529,000 - 10,262,000
Work in a fully remote environment
Opportunity to work on cutting-edge AI projects
Remote Data Scientist - 60713
Remote Data Scientist - 60713

Turing • India

Remote
INR 5,273,000 - 7,910,000
Fully remote environment
Work on cutting-edge AI projects
Flexible working hours
Remote Data Engineer - 60713
Remote Data Engineer - 60713

Turing • Dadri

Remote
INR 1,200,000 - 1,800,000
Work in a fully remote environment
Engage in cutting-edge AI projects
Flexible working hours with proper overlap
Remote Data Engineer - 60713
Remote Data Engineer - 60713

Turing • Hyderabad

Remote
INR 5,273,000 - 9,228,000
Fully remote work
Cutting-edge AI projects
Flexible hours
+1
Remote Data Engineer - 60713
Remote Data Engineer - 60713

Turing • Hyderabad

Remote
INR 1,200,000 - 2,400,000
Fully remote work
Cutting-edge AI projects
Flexible work hours
+1
Remote Data Engineer - 60713
Remote Data Engineer - 60713

Turing • Bengaluru

Remote
Fully remote environment
Opportunity to work on cutting-edge AI projects