Senior Software Engineer, AI Training - UK

United States Digital Space LLC

United Kingdom

Remote

GBP 104,000 - 208,000

Part time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a senior-level software engineer to evaluate AI coding agents. This is not traditional coding work; you will assess how models think and respond like engineers.

You will provide detailed, opinionated feedback on usefulness, explanations, and engineering judgment, helping define what great looks like when using tools like Cursor, Codex and Claude Code. Remote, worldwide with 10–20 hours per week.

Qualifications

  • Senior/Staff/Principal-level engineer with strong backend/frontend expertise.
  • Fluent English (B2 or higher) for feedback and video explanations.
  • Hands-on experience with OpenAI Codex, Claude Code or Cursor.

Responsibilities

  • Evaluate AI-generated coding interactions end to end.
  • Judge usefulness, high-level accuracy, and engineering judgment.
  • Assess explanations and reasoning, not just the code.
  • Differentiate what makes a response quality level (2 vs 4).
  • Provide clear, opinionated written feedback on what worked and what didn’t.
  • Help define what great looks like with tools like Cursor, Codex and Claude Code.

Skills

TypeScript/JavaScript
Python
English proficiency

Tools

OpenAI Codex
Claude Code
Cursor

Job description

Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week

We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.

This is not a traditional engineering role. You won't be writing production code. You'll be evaluating something harder: whether the model thinks like a great engineer.

What this role actually is

You will assess how AI coding agents behave in real-world scenarios, focusing on:

  • Whether the response makes sense
  • Whether the preamble and reasoning are useful
  • Whether the output reflects strong engineering judgment
  • Whether the interaction feels right to an experienced developer

This role is about engineering taste. Syntax correctness is the easy part.

What you'll be doing
  • Evaluate AI-generated coding interactions end to end
  • Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning, not just the code
  • Distinguish between levels of response quality (what makes something a 2 vs a 4)
  • Give clear, opinionated feedback in writing on what worked, what didn't, and what felt off or misleading
  • Help define what great looks like when working with tools like Cursor, Codex and Claude Code
What we mean by "taste"

We're looking for engineers who can answer questions like:

  • Does this feel like something a strong engineer would actually say?
  • Is this explanation helpful, or just technically correct?
  • Is the model guiding the user well, or just dumping output?
  • Would this interaction build or erode trust?

You should be comfortable making subjective but rigorous judgments, and explaining them clearly.

Who you are
  • Senior, Staff or Principal-level engineer (or equivalent experience)
  • Strong background in TypeScript/JavaScript or Python
  • Hands‑on experience with at least one of OpenAI Codex, Claude Code or Cursor
  • Deep familiarity with modern AI-assisted dev workflows
  • Able to evaluate code without executing it or reviewing every line
  • Strong written and spoken English (B2 or above). You'll be writing detailed feedback and recording short video explanations, so this is a hard requirement
  • Comfortable giving direct, opinionated feedback
  • High bar for what good engineering looks like
Nice to have
  • Prior exposure to prompt design or evaluation workflows
  • Experience mentoring senior engineers or defining engineering standards
Engagement details
  • Rate: $100-$200/hour depending on experience and location
  • Hours: 10-20 hours/week, flexible
  • Duration: ongoing. Projects run from about two weeks to a few months each, and we offer new ones to evaluators who do well
  • Start: as soon as you clear the take-home and a project has an open seat
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Engineering
Engineering

Work From Anywhere • United Kingdom

Remote
GBP 104,000 - 208,000
Senior AI Coding Evaluator (Remote) — Engineering Taste
Senior AI Coding Evaluator (Remote) — Engineering Taste

Work From Anywhere • United Kingdom

Remote
GBP 104,000 - 208,000
Senior Software Engineer — AI Coding Evaluator & Feedback
Senior Software Engineer — AI Coding Evaluator & Feedback

United States Digital Space LLC • United Kingdom

Remote
GBP 104,000 - 208,000
Remote Software Engineer
Remote Software Engineer

turing • United Kingdom

Remote
GBP 42,000 - 99,000
AI Engineer
AI Engineer

RemoteJobsOne • Greater London

Remote
GBP 62,000 - 125,000
Senior Software Engineer – AI Training Projects
Senior Software Engineer – AI Training Projects

Andrews Recruitment Group • Birmingham

Remote
GBP 31,000 - 154,000
Fully remote
Flexible hours
Senior Software Engineer – AI Training Projects
Senior Software Engineer – AI Training Projects

Andrews Recruitment Group • Greater London

Remote
GBP 31,000 - 154,000
Java Developer
Java Developer

Turing • Greater London

On-site
GBP 55,000 - 83,000
Senior Software Engineer – AI Training Projects
Senior Software Engineer – AI Training Projects

Andrews Recruitment Group • Glasgow

Remote
GBP 31,000 - 154,000
Senior Software Engineer – AI Training Projects
Senior Software Engineer – AI Training Projects

Andrews Recruitment Group • Manchester

Remote
GBP 31,000 - 154,000
Fully remote
Flexible hours
Contractor position
+1