Turn this role into an interview — a resume and cover letter built around what this employer wants.
United States Digital Space LLC is seeking a senior-level evaluator to critically assess AI-assisted coding interactions in real-world scenarios. You will judge usefulness, accuracy, and engineering judgment while avoiding execution of code.
Feedback will be written and shared through Loom-style walkthroughs. Ideal candidates have strong TypeScript/JavaScript or Python skills, experience with Codex/Claude/Cursor, and excellent English communication.
Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week
Check out this Loom video for more details: https://www.loom.com/share/b0d1b0bf24c44ae8b95dca84b9db60e5
We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
This is not a traditional engineering role. You won't be writing production code. You'll be evaluating something harder: whether the model thinks like a great engineer.
You will assess how AI coding agents behave in real-world scenarios, focusing on:
This role is about engineering taste. Syntax correctness is the easy part.
Evaluate AI-generated coding interactions end to end
We're looking for engineers who can answer questions like:
You should be comfortable making subjective but rigorous judgments, and explaining them clearly.
Senior, Staff or Principal-level engineer (or equivalent experience)
Prior exposure to prompt design or evaluation workflows
Rate: $100-$200/hour depending on experience and location