Senior Software Engineer - AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)

G2i Inc.

Turkey

On-site

TRY 3,286,000 - 13,145,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

G2i Inc. is seeking a Senior AI Interaction Evaluator (Codex / Claude Code) contract role. You will assess AI coding agents on realism, usefulness, and engineering judgment without writing production code.

Start ASAP, 10–20 hrs/week, hourly pay between $50 and $200. You’ll judge how the model thinks like a great engineer and help define high-quality interaction standards for tools like Cursor.

Qualifications

  • Strong background in Python programming.
  • Experience with OpenAI Codex or similar coding assistants.
  • Familiarity with Cursor and AI-first IDEs.
  • Ability to judge quality of engineering explanations.
  • Comfort giving direct, opinionated feedback.

Responsibilities

  • Evaluate AI-generated coding interactions end-to-end.
  • Judge usefulness, correctness, and engineering judgment at a high level.
  • Assess quality of explanations and reasoning.
  • Provide clear, opinionated feedback on what worked and what didn’t.
  • Define what great looks like when interacting with Cursor.

Skills

Python programming
Engineering judgment
Feedback writing

Tools

OpenAI Codex
Cursor

Job description

Senior AI Interaction Evaluator (Codex / Claude Code)

Contract | $50-200/hr | 10–20 hrs/week | Start ASAP (through early May)

Check out this Loom video for more details!

We’re looking for a highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.

This is not a traditional engineering role.

You won’t be writing production code.

You’ll be evaluating something harder: whether the model thinks like a great engineer.

What This Role Actually Is

You will assess how AI coding agents behave in real-world scenarios — focusing on:

  • Whether the response makes sense
  • Whether the preamble and reasoning are useful
  • Whether the output reflects strong engineering judgment
  • Whether the interaction feels right to an experienced developer

This role is about engineering taste — not syntax correctness.

What You’ll Be Doing

  • Evaluate AI-generated coding interactions end-to-end
  • Judge whether outputs are:
  • Useful
  • Correct (at a high level)
  • Aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning, not just code
  • Distinguish between different levels of response quality (e.g. what makes something a 2 vs 4)
  • Provide clear, opinionated feedback on:
  • What worked
  • What didn’t
  • What felt “off” or misleading
  • Help define what great looks like when interacting with tools like Cursor

What We Mean by “Taste”

We’re specifically looking for engineers who can answer questions like:

  • Does this feel like something a strong engineer would actually say?
  • Is this explanation helpful, or just technically correct?
  • Is the model guiding the user well, or just dumping output?
  • Would this interaction build or erode trust?

You should be comfortable making subjective but rigorous judgments.

Who You Are

  • Strong background in one of the below:
  • Python
  • Hands-on experience using:
  • OpenAI Codex
  • Cursor
  • Deep familiarity with modern AI-assisted dev workflows
  • Able to evaluate code without needing to fully execute or deeply review every line
  • Comfortable giving direct, opinionated feedback
  • High bar for what “good engineering” looks like

Nice to Have

  • Experience with tools like Cursor or similar AI-first IDEs
  • Prior exposure to prompt design or evaluation workflows
  • Experience mentoring senior engineers or defining engineering standards
  • US and Canada up to $200/hr
  • EU and Latam up to $150/hr
  • Other locations up to $100/hr
  • Hours: ~10–20 hours/week
  • Duration: Through early May (with possible extension)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Interaction Evaluator - Engineering Taste
Senior AI Interaction Evaluator - Engineering Taste

G2i Inc. • Turkey

On-site
TRY 3,286,000 - 13,145,000
Forward-Deployed AI Engineer
Forward-Deployed AI Engineer

Zaigo • Fatih

On-site
TRY 600,000 - 1,000,000
C++ Developer - Remote
C++ Developer - Remote

YO IT Consulting • Fatih

On-site
TRY 80,000 - 100,000
Lead AI Engineer with Machine Learning
Lead AI Engineer with Machine Learning

EPAM Systems • Turkey

On-site
TRY 5,747,000 - 8,621,000
Continuous upskilling
Private health insurance
Learning platforms access
+1
Content Writer - Remote
Content Writer - Remote

YO IT Consulting • Fatih

On-site
TRY 1,583,448 - 3,166,896
AI Trainer Lead - Remote
AI Trainer Lead - Remote

YO IT Consulting • Fatih

On-site
TRY 1,292,307 - 2,261,538
Generative AI Operations Engineer (GenAI Ops)
Generative AI Operations Engineer (GenAI Ops)

EPAM Systems • Turkey

On-site
TRY 2,874,000 - 4,789,000
Continuous upskilling
Private health insurance
English courses
+3
Robotics Evaluation Specialist - Remote
Robotics Evaluation Specialist - Remote

YO IT Consulting • Fatih

On-site
TRY 2,099,717 - 3,674,506
Senior Data Engineer, AI
Senior Data Engineer, AI

EPAM Systems • Turkey

On-site
TRY 5,744,000 - 8,617,000
Private health insurance
Learning & development programs
English courses
+1
Lead AI Engineer
Lead AI Engineer

EPAM Systems • Turkey

On-site
TRY 5,747,000 - 8,621,000
Private health insurance
English courses
Career development