Engineering

Work From Anywhere

United Kingdom

Remote

GBP 104,000 - 208,000

Part time

13 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Work From Anywhere is seeking a highly experienced software engineer (SR+) to evaluate AI coding agents such as OpenAI Codex and Claude Code. This is not a traditional coding role; you will judge whether the model thinks like a great engineer and if responses are useful and well-reasoned.

The position is remote worldwide, 10-20 hours per week, with take-home tasks and Loom walkthroughs. Applicants should be senior-level, hands-on with Codex/Claude Code/Cursor, and able to produce clear written

Qualifications

  • Senior, Staff or Principal-level software engineer.
  • Hands-on with Codex, Claude Code or Cursor.
  • Strong written English and ability to provide clear feedback.
  • Experience with AI-assisted dev workflows.

Responsibilities

  • Evaluate AI-generated coding interactions end to end.
  • Provide clear, opinionated feedback on usefulness and accuracy.
  • Assess explanations and reasoning beyond code.
  • Help define criteria of excellence with tools like Codex/Cursor.

Skills

Engineering judgment
Feedback writing
English proficiency
Subjective evaluation
Video explanations
High standards

Tools

OpenAI Codex
Claude Code
Cursor

Job description

Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week

Check out this Loom video for more details:

We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.

This is not a traditional engineering role. You won't be writing production code. You'll be evaluating something harder: whether the model thinks like a great engineer.

What this role actually is

You will assess how AI coding agents behave in real-world scenarios, focusing on:

Whether the response makes sense
  • Whether the preamble and reasoning are useful
  • Whether the output reflects strong engineering judgment
  • Whether the interaction feels right to an experienced developer

This role is about engineering taste. Syntax correctness is the easy part.

What you'll be doing
  • Evaluate AI-generated coding interactions end to end
  • Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning, not just the code
  • Distinguish between levels of response quality (what makes something a 2 vs a 4)
  • Give clear, opinionated feedback in writing on what worked, what didn't, and what felt off or misleading
  • Help define what great looks like when working with tools like Cursor, Codex and Claude Code
What we mean by "taste"

We're looking for engineers who can answer questions like:

  • Does this feel like something a strong engineer would actually say?
  • Is this explanation helpful, or just technically correct?
  • Is the model guiding the user well, or just dumping output?
  • Would this interaction build or erode trust?

You should be comfortable making subjective but rigorous judgments, and explaining them clearly.

Who you are
  • Senior, Staff or Principal-level engineer (or equivalent experience)
  • Hands-on experience with at least one of OpenAI Codex, Claude Code or Cursor
  • Deep familiarity with modern AI-assisted dev workflows
  • Able to evaluate code without executing it or reviewing every line
  • Strong written and spoken English (B2 or above). You'll be writing detailed feedback and recording short video explanations, so this is a hard requirement
  • Comfortable giving direct, opinionated feedback
  • High bar for what good engineering looks like
Nice to have
  • Prior exposure to prompt design or evaluation workflows
  • Experience mentoring senior engineers or defining engineering standards
  • Rate: $100-$200/hour depending on experience and location
  • Duration: ongoing. Projects run from about two weeks to a few months each, and we offer new ones to evaluators who do well
  • Start: as soon as you clear the take-home and a project has an open seat
  • Process: one take-home evaluation exercise with a recorded Loom walkthrough. No interview

Originally posted on Himalayas

Ready?

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, AI Training - UK
Senior Software Engineer, AI Training - UK

United States Digital Space LLC • United Kingdom

Remote
GBP 104,000 - 208,000
Senior AI Coding Evaluator (Remote) — Engineering Taste
Senior AI Coding Evaluator (Remote) — Engineering Taste

Work From Anywhere • United Kingdom

Remote
GBP 104,000 - 208,000
Remote Software Engineer
Remote Software Engineer

turing • United Kingdom

Remote
GBP 42,000 - 99,000
Software Engineer - Fully Remote | Up to $150/hr
Software Engineer - Fully Remote | Up to $150/hr

Obsidian • Greater London

Remote
GBP 55,000 - 83,000
Senior Software Engineer — AI Coding Evaluator & Feedback
Senior Software Engineer — AI Coding Evaluator & Feedback

United States Digital Space LLC • United Kingdom

Remote
GBP 104,000 - 208,000
Rust Software Engineer — AI Data & Evaluation (Contract)
Rust Software Engineer — AI Data & Evaluation (Contract)

Turing • Greater London

On-site
GBP 51,000 - 102,000
Remote Senior Software Engineer (LLM) - 34953
Remote Senior Software Engineer (LLM) - 34953

Turing • London

On-site
GBP 62,504 - 125,008
Flexible working hours
Collaborate with top researchers
Remote Software Engineer (UK)
Remote Software Engineer (UK)

Turing • Oxford

On-site
GBP 59,023 - 88,534
Java Developer
Java Developer

Turing • Greater London

On-site
GBP 55,000 - 83,000
Software Engineering Evaluator - Expert
Software Engineering Evaluator - Expert

Mercor • Greater London

Remote
GBP 55,000 - 117,000