Software Engineer - Evaluation Harness & Code Review

Terac

United States

Remote

USD 76,000 - 103,000

Full time

39 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Terac is conducting a remote study to benchmark coding environments used to evaluate AI agents. You will review realistic programming tasks and their evaluation harnesses to ensure they reflect real-world software engineering challenges.

During guided sessions you’ll explain your thought process aloud, participate in screen-sharing exercises, and provide feedback on logic, test cases, and overall environment design. Compensation is $65 per hour.

Qualifications

  • Active professional experience as a software engineer.
  • Hands-on experience building or maintaining test suites and evaluation harnesses.
  • Comfortable reviewing code and explaining technical concepts aloud.
  • Familiarity with complex real-world software architecture.

Responsibilities

  • Review realistic programming tasks for accuracy and difficulty.
  • Verify the logic and test cases within provided evaluation harnesses.
  • Assess whether coding environments effectively measure software engineering skills.
  • Walk us through your thought process while analyzing complex code structures.

Skills

Software engineering
Test suites
Evaluation harnesses
Code reviews
Communication

Job description

Terac is conducting a remote study to benchmark coding environments used to evaluate AI agents. You will review realistic programming tasks and their evaluation harnesses to ensure they reflect real-world software engineering challenges.

During guided sessions you’ll explain your thought process aloud, participate in screen-sharing exercises, and provide feedback on logic, test cases, and overall environment design. Compensation is $65 per hour.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineers: Paid Code Review for AI Agent Evaluation
Software Engineers: Paid Code Review for AI Agent Evaluation

Terac • United States

Remote
USD 76,000 - 103,000
AI Evaluation Engineer — Review Coding Tasks
AI Evaluation Engineer — Review Coding Tasks

Remote Jobs • United States

Remote
USD 114,341,000 - 137,760,000
LATAM Software Engineers: Coding Tasks for AI Evaluation
LATAM Software Engineers: Coding Tasks for AI Evaluation

Remote Jobs • United States

Remote
USD 114,341,000 - 137,760,000
South Asian Software Engineers: Coding Tasks for AI Evaluation
South Asian Software Engineers: Coding Tasks for AI Evaluation

AI Trainer Jobs • United States

Remote
USD 70,000 - 76,000
Remote Senior Engineer — AI Coding Agent Evaluation
Remote Senior Engineer — AI Coding Agent Evaluation

Remote Worker LTD. • United States

Remote
USD 138,000 - 276,000
Remote South Asia Software Engineers - Evaluation Study
Remote South Asia Software Engineers - Evaluation Study

AI Trainer Jobs • United States

Remote
USD 70,000 - 76,000
Remote Software Engineer for AI Code Evaluation
Remote Software Engineer for AI Code Evaluation

OpenTrain AI • Northern (KY)

Hybrid
USD 138,000 - 207,000
Remote contractor
Part-time engagement
Hourly pay $100–$150 USD
+1
Senior AI Coding Evaluation Engineer - Remote, 10-20h/wk
Senior AI Coding Evaluation Engineer - Remote, 10-20h/wk

Remote Jobs • United States

Remote
USD 138,000 - 276,000
Remote Software Engineer - AI Coding Tools Study Participant
Remote Software Engineer - AI Coding Tools Study Participant

AI Trainer Jobs • United States

Remote
USD 57 - 69
Remote Java Engineer for AI Code Evaluation
Remote Java Engineer for AI Code Evaluation

turing • San Francisco (CA)

Remote
USD 14,000 - 55,000