Software Engineers: Paid Code Review for AI Agent Evaluation

Terac

United States

Remote

USD 138,000 - 413,000

Part time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote study participation
Compensation for time

Job summary

Terac is conducting a remote, paid study to evaluate coding environments and evaluation harnesses used to benchmark AI agents. Participants will review tasks and harnesses and provide feedback on realism and usefulness.

We are seeking practicing software engineers with hands-on experience building and testing complex systems, including writing test suites, evaluation harnesses, or performing comprehensive code reviews. Sessions are conducted via guided conversations and screen-sharing.

Qualifications

  • Practicing software engineers with experience building and testing complex systems.
  • Direct experience writing test suites, evaluation harnesses, or performing code reviews.
  • Background in backend, full-stack, ML engineering, or software architecture is preferred.

Responsibilities

  • Review coding tasks and evaluation harnesses for realism.
  • Verify logic, test cases, and structure of the environments provided.
  • Provide technical feedback via guided conversation and screen-sharing.

Skills

Software engineering
Test suite development
Code reviews
Evaluation harnesses
Backend development

Education

Bachelor's degree in Computer Science or related field

Tools

Git
CI/CD

Job description

What We're Researching

We're running a paid study on the coding environments and programming tasks used to benchmark artificial intelligence agents. Creating robust evaluation harnesses ensures that AI models are tested against realistic software engineering scenarios. This work directly feeds into improving how autonomous agents handle complex coding objectives.

How It Works

During this remote session, you will review a series of coding tasks and their corresponding evaluation harnesses. You will assess whether the programming challenges accurately reflect real-world software engineering problems. We will ask you to verify the logic, test cases, and overall structure of the environments provided. Your technical feedback will be captured through a guided conversation and screen-sharing exercises.

Who This Is For

We are looking for practicing software engineers with strong backgrounds in building and testing complex systems. Candidates should have direct experience writing test suites, evaluation harnesses, or comprehensive code reviews. We welcome backend engineers, full-stack developers, machine learning engineers, and software architects.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineers: Paid Code Review for AI Agent Evaluation
Software Engineers: Paid Code Review for AI Agent Evaluation

AI Trainer Jobs • United States

Remote
USD 83,000 - 96,000
AI Evaluation Engineer — Review Coding Tasks
AI Evaluation Engineer — Review Coding Tasks

Remote Jobs • United States

Remote
USD 114,341,000 - 137,760,000
Remote Software Engineer - Evaluation Harness Reviewer
Remote Software Engineer - Evaluation Harness Reviewer

AI Trainer Jobs • United States

Remote
USD 83,000 - 96,000
South Asian Software Engineers: Coding Tasks for AI Evaluation
South Asian Software Engineers: Coding Tasks for AI Evaluation

AI Trainer Jobs • United States

Remote
USD 70,000 - 76,000
Remote AI Benchmarking & Code Evaluation Engineer
Remote AI Benchmarking & Code Evaluation Engineer

Terac • United States

Remote
USD 138,000 - 413,000
Remote study participation
Compensation for time
LATAM Software Engineers: Coding Tasks for AI Evaluation
LATAM Software Engineers: Coding Tasks for AI Evaluation

Remote Jobs • United States

Remote
USD 114,341,000 - 137,760,000
Software Engineers: Feedback on AI Coding Tools
Software Engineers: Feedback on AI Coding Tools

AI Trainer Jobs • United States

Remote
USD 57 - 69
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • United States

On-site
USD 41,000 - 69,000
Senior Software Engineer — AI Coding Evaluator
Senior Software Engineer — AI Coding Evaluator

AI Trainer Jobs • United States

Remote
USD 83,000 - 124,000
Site Reliability Engineering AI Evaluator
Site Reliability Engineering AI Evaluator

AI Trainer Jobs • United States

Remote
USD 83,000 - 165,000
Remote work
Contractor position
Flexible hours