AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)

Triwill Group

United States

On-site

USD 138,000 - 276,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work
Flexible schedule
Short-term engagement
Take-home evaluation

Job summary

Jobgether is seeking a senior engineer for a remote, contract-based role to evaluate AI coding agents in real-world software-development scenarios.

The work focuses on reasoning, explanations, and developer guidance rather than just syntactically correct code, with a flexible 10–20 hour weekly commitment and hourly compensation.

Qualifications

  • Senior engineering background with substantial practical software engineering experience.
  • Strong professional background in Python and/or TypeScript/JavaScript.
  • Hands-on experience with AI coding tools such as OpenAI Codex, Claude Code, Cursor.
  • Ability to articulate clear, concise, well-supported opinions about AI-generated interactions.
  • Maintain high standards for software engineering, developer experience, and technical decision-making.

Responsibilities

  • Evaluate AI interactions end-to-end for usefulness and high-level accuracy.
  • Assess engineering judgment and practical decision-making in AI responses.
  • Review explanations, reasoning, and guidance to aid developers.
  • Measure response quality and identify strengths, weaknesses, and misleading aspects.
  • Provide actionable, direct feedback on what works and what doesn’t.

Skills

Python
TypeScript/JavaScript
AI coding experience
Communication skills
Engineering judgment
High quality standards

Tools

OpenAI Codex
Claude Code
Cursor

Job description

This is a senior-level contract opportunity focused on evaluating how modern AI coding agents interact with experienced software engineers. Rather than developing production software, you’ll use your engineering expertise to assess the quality, usefulness, and judgment demonstrated by AI-generated responses. You’ll evaluate tools such as Codex, Claude Code, and Cursor across realistic software-development scenarios. The role goes beyond checking whether code is technically correct, focusing on reasoning, explanations, developer guidance, and overall interaction quality. You’ll apply rigorous engineering judgment to distinguish genuinely strong AI assistance from responses that are merely plausible or syntactically correct. Your feedback will help establish clearer standards for what excellent AI-assisted development should look like. The engagement is fully remote, flexible, and designed for experienced engineers who want to contribute directly to the evolution of AI coding tools.

Accountabilities
  • Evaluate AI interactions: Review AI-generated coding interactions end to end and determine whether responses are useful, accurate at a high level, and consistent with strong engineering practices.
  • Assess engineering judgment: Evaluate whether coding agents demonstrate sound technical reasoning, appropriate decision-making, and practical engineering judgment rather than simply producing working-looking code.
  • Review explanations and reasoning: Assess the quality of preambles, explanations, reasoning, and guidance, identifying whether they genuinely help developers understand and solve problems.
  • Measure response quality: Distinguish between different levels of AI response quality and identify the characteristics that separate adequate interactions from exceptional ones.
  • Provide actionable feedback: Deliver clear, direct, and opinionated assessments covering what worked, what failed, and what felt misleading, ineffective, or inconsistent with experienced engineering practice.
  • Evaluate developer experience: Consider whether an interaction would build trust with an experienced developer, provide useful direction, or instead create confusion and unnecessary work.
  • Shape evaluation standards: Help define what "great" looks like when developers work with AI coding agents and AI-first development environments.
  • Apply independent expertise: Make subjective but rigorous judgments without needing to execute or deeply inspect every line of generated code.
Requirements
  • Senior engineering background: Staff-, Principal-, or similarly experienced software engineer, or equivalent depth of practical engineering experience.
  • Programming expertise: Strong professional background in Python and/or TypeScript/JavaScript, with the ability to quickly understand different software-development approaches.
  • AI coding experience: Hands-on experience with modern AI coding tools such as OpenAI Codex, Claude Code, Cursor, or comparable AI-assisted development platforms.
  • AI-assisted development knowledge: Deep familiarity with contemporary AI-enabled software-engineering workflows and how developers collaborate with coding agents.
  • Strong engineering taste: Ability to recognize whether a response reflects the thinking, communication style, trade-offs, and technical standards expected from an excellent engineer.
  • Critical evaluation ability: Comfortable assessing code and technical reasoning at a high level without needing to execute or manually review every implementation detail.
  • Communication skills: Able to articulate clear, concise, and well-supported opinions about the strengths and weaknesses of AI-generated interactions.
  • High quality standards: A consistently high bar for software engineering, developer experience, clarity, and technical decision-making.
  • Preferred experience: Exposure to AI-first IDEs, prompt design, model evaluation, or structured AI evaluation workflows is advantageous.
  • Additional advantage: Experience mentoring senior engineers or establishing engineering standards and best practices is a plus.
Benefits
  • Compensation: $100–$200 per hour.
  • Flexible schedule: Approximately 10–20 hours per week, allowing you to balance the engagement with other commitments.
  • Remote work: Fully remote contract opportunity across eligible locations.
  • Short-term engagement: Initial engagement through early May, with the possibility of extension.
  • Fast start: Opportunity to begin as soon as possible.
  • High-impact work: Help shape how AI coding agents are evaluated and improve the quality of AI-assisted software development.
  • Senior-level contribution: Apply years of engineering experience to challenging questions around AI reasoning, developer experience, and engineering quality.
  • Streamlined selection process: Take-home evaluation exercise followed by one behavioral interview.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)
AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)

G2i Inc. • Arlington (WA)

On-site
USD 138,000 - 276,000
Senior Python Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)
Senior Python Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2i Inc. • Miami (FL)

On-site
USD 138,000 - 276,000
Senior Rust Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)
Senior Rust Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)

G2i Inc. • Miami (FL)

On-site
USD 138,000 - 276,000
Senior Engineer - AI Interaction Evaluator - US and Canada only
Senior Engineer - AI Interaction Evaluator - US and Canada only

G2i Inc. • Miami (FL)

On-site
USD 138,000 - 276,000
Competitive pay rate
Flexible schedule
Possible extension
+1
AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)
AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)

G2i Inc. • Miami (FL)

On-site
USD 137,760 - 275,520
Python Engineer, AI Code Reviewer
Python Engineer, AI Code Reviewer

Far Coder • Northern (KY)

Hybrid
USD 138,000 - 276,000
Senior AI Interaction Evaluator for Codex & Claude Code
Senior AI Interaction Evaluator for Codex & Claude Code

Doist • Miami (FL)

On-site
USD 138,000 - 276,000
Senior AI Interaction Evaluator
Senior AI Interaction Evaluator

G2i Inc. • United States

On-site
USD 83,000 - 124,000
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Alabama

Remote
Flexible working hours
Competitive pay up to $80/hour
Experience in advanced AI projects
Software Developer, AI Coding Evaluation
Software Developer, AI Coding Evaluation

OpenTrain AI • Northern (KY)

Hybrid
USD 83,000 - 165,000