Python Engineer, AI Coding Agent Evaluator

g2i

Deutschland

Remote

EUR 119.000 - 238.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Mach aus dieser Rolle ein Bewerbungsgespräch — ein Lebenslauf und ein Anschreiben, die genau auf das zugeschnitten sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

g2i seeks a Senior AI Interaction Evaluator to assess AI coding interactions from Codex and Claude Code. You will judge usefulness, reasoning quality, and alignment with engineer thinking, without writing production code.

Expected to provide direct, opinionated feedback and help define what great looks like when using Cursor in AI-assisted development workflows.

Qualifikationen

  • Staff / Principal-level engineer (or equivalent experience).
  • Strong background in TypeScript/JavaScript or Python.
  • Hands-on experience using OpenAI Codex, Claude Code and Cursor.
  • Deep familiarity with modern AI-assisted dev workflows.
  • Able to evaluate code without needing to fully execute.

Aufgaben

  • Evaluate AI-generated coding interactions end-to-end.
  • Assess usefulness, correctness at a high level and engineering judgment.
  • Provide clear, opinionated feedback on what worked and what didn't.
  • Help define what great looks like when interacting with Cursor.

Kenntnisse

TypeScript/JavaScript
Python
OpenAI Codex
Claude Code
Cursor
AI-assisted dev workflows

Tools

OpenAI Codex
Claude Code
Cursor

Jobbeschreibung

Senior AI Interaction Evaluator (Codex / Claude Code)

Contract | $100-$200/hour | 10-20 hrs/week | Start ASAP (through early May)

Check out this Loom video for more details!

We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role. You won't be writing production code. You'll be evaluating something harder: whether the model thinks like a great engineer.

What This Role Actually Is
  • Whether the response makes sense
  • Whether the preamble and reasoning are useful
  • Whether the output reflects strong engineering judgment
  • Whether the interaction feels right to an experienced developer
What You'll Be Doing

Evaluate AI-generated coding interactions end-to-end

  • Useful
  • Correct (at a high level)
  • Aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning , not just code
  • Distinguish between different levels of response quality (e.g. what makes something a 2 vs 4 )

Provide clear, opinionated feedback on:

  • What worked
  • What didn't
  • What felt "off" or misleading
  • Help define what great looks like when interacting with tools like Cursor
What We Mean by "Taste"
  • Does this feel like something a strong engineer would actually say?
  • Is this explanation helpful, or just technically correct?
  • Is the model guiding the user well, or just dumping output?
  • Would this interaction build or erode trust?
Who You Are
  • Staff / Principal-level engineer (or equivalent experience)
  • Strong background in one of the below:
  • TypeScript / JavaScript
  • Python
  • Hands-on experience using:
  • OpenAI Codex
  • Claude Code
  • Cursor
  • Deep familiarity with modern AI-assisted dev workflows
  • Able to evaluate code without needing to fully execute or deeply review every line
  • Comfortable giving direct, opinionated feedback
  • High bar for what "good engineering" looks like
Nice to Have
  • Experience with tools like Cursor or similar AI-first IDEs
  • Prior exposure to prompt design or evaluation workflows
  • Experience mentoring senior engineers or defining engineering standards
Engagement Details
  • Rate: $100-$200/hour
  • Hours: ~10-20 hours/week
  • Duration: Through early May (with possible extension)
  • Start: ASAP
Process:
  • Take-home evaluation exercise
  • One behavioral interview
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Python Engineer - LLM Code Evaluation
Senior Python Engineer - LLM Code Evaluation

g2i • Deutschland

Vor Ort
EUR 119.000 - 238.000
TypeScript Engineer, AI Coding Agent Evaluator
TypeScript Engineer, AI Coding Agent Evaluator

g2i • Deutschland

Vor Ort
EUR 119.000 - 238.000
Senior Software Engineer - Agent Evaluation
Senior Software Engineer - Agent Evaluation

aitrainer • Deutschland

Vor Ort
EUR 47.779 - 71.669
Remote Senior Software Engineer (LLM) - 34953
Remote Senior Software Engineer (LLM) - 34953

Turing • Deutschland

Vor Ort
EUR 47.303 - 118.258
Senior AI Agent Evaluation Engineer
Senior AI Agent Evaluation Engineer

aitrainer • Deutschland

Vor Ort
EUR 34.453 - 60.293
Applied AI Engineer, Codex
Applied AI Engineer, Codex

AI Chopping Block, Inc. • München

Hybrid
EUR 163.000 - 221.000
Applied AI Engineer, Codex
Applied AI Engineer, Codex

OpenAI • München

Vor Ort
EUR 140.000 - 210.000
Relocation assistance
Hybrid work model
AI Engineer, Quality & Evals (m/f/x)
AI Engineer, Quality & Evals (m/f/x)

Cortea AI • Berlin

Vor Ort
EUR 90.000 - 140.000
Equity
Flexible vacation
Team lunches
+2
DevOps / SRE / Cloud Engineer (Coding Agent Experience)
DevOps / SRE / Cloud Engineer (Coding Agent Experience)

aitrainer • Deutschland

Vor Ort
EUR 43.066 - 68.906
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Berlin

Vor Ort
EUR 50.711 - 72.224
Competitive hourly rates
Flexible schedule
Experience in advanced AI projects