AI Evaluation Engineer — Review Coding Tasks

Remote Jobs

United States

Remote

USD 114,341,000 - 137,760,000

Full time

41 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Terac is conducting a paid study to evaluate coding tasks used for AI agent evaluation. Our team is building a suite of programming environments to test complex software capabilities and ensure frameworks are realistic, accurate, and well calibrated.

We are seeking software engineers in the United States with experience building, reviewing, or testing evaluation harnesses and coding tasks. Candidates should be comfortable discussing technical architectures and providing actionable feedback on

Qualifications

  • Professional software engineer with strong technical background.
  • Experience building or reviewing evaluation harnesses and coding tasks.
  • Comfortable discussing technical architectures and evaluation frameworks.

Responsibilities

  • Review realistic programming tasks designed for AI evaluation.
  • Assess the accuracy and structure of various evaluation harnesses.
  • Walk us through your thought process while checking code validity.
  • Provide actionable feedback on how to improve task complexity and realism.

Skills

Software engineer
Eval harness familiarity

Job description

Terac is conducting a paid study to evaluate coding tasks used for AI agent evaluation. Our team is building a suite of programming environments to test complex software capabilities and ensure frameworks are realistic, accurate, and well calibrated.

We are seeking software engineers in the United States with experience building, reviewing, or testing evaluation harnesses and coding tasks. Candidates should be comfortable discussing technical architectures and providing actionable feedback on

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer - Evaluation Harness & Code Review
Software Engineer - Evaluation Harness & Code Review

Terac • United States

Remote
USD 76,000 - 103,000
Software Engineers: Paid Code Review for AI Agent Evaluation
Software Engineers: Paid Code Review for AI Agent Evaluation

Terac • United States

Remote
USD 76,000 - 103,000
LATAM Software Engineers: Coding Tasks for AI Evaluation
LATAM Software Engineers: Coding Tasks for AI Evaluation

Remote Jobs • United States

Remote
USD 114,341,000 - 137,760,000
South Asian Software Engineers: Coding Tasks for AI Evaluation
South Asian Software Engineers: Coding Tasks for AI Evaluation

AI Trainer Jobs • United States

Remote
USD 70,000 - 76,000
Remote AI Evaluation Engineer for Code Tasks
Remote AI Evaluation Engineer for Code Tasks

YO AI Labs • California (MO)

Remote
USD 50,000 - 80,000
AI Evaluation & Task Design Engineer
AI Evaluation & Task Design Engineer

Biz Tech Consultants • United States

Remote
USD 120,000 - 180,000
Senior AI Code Evaluator (Python/TypeScript)
Senior AI Code Evaluator (Python/TypeScript)

Turing • United States

Remote
USD 83,000 - 152,000
Remote South Asia Software Engineers - Evaluation Study
Remote South Asia Software Engineers - Evaluation Study

AI Trainer Jobs • United States

Remote
USD 70,000 - 76,000
AI Code Evaluator (JavaScript) - RLHF Tasks, Contract
AI Code Evaluator (JavaScript) - RLHF Tasks, Contract

Biz Tech Consultants • United States

Remote
USD 83,000 - 152,000
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • United States

On-site
USD 41,000 - 69,000