AI Evaluation Engineer: Coding Task Architect

United States Digital Space LLC

United States

Remote

USD 55,000 - 69,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is building a dataset to evaluate AI coding agents by creating realistic developer environments with codebases, infrastructure, tickets, docs, and a believable development history. You will design tasks from intermediate states and craft what a solution must achieve to be considered solved.

The role emphasizes writing tests that accept all valid approaches while rejecting incorrect ones, and iterating on tasks based on QA feedback to keep evaluations fair and

Qualifications

  • 5+ years in software development.
  • Proficiency with Python (FastAPI) and JavaScript/TypeScript (React).
  • Experience with Docker, PostgreSQL, Kafka, and Redis.
  • Experience writing tests (functional, integration).
  • English proficiency - B2+.

Responsibilities

  • Build realistic developer environments—a virtual company with codebase, infra, tickets and docs.
  • Design tasks from intermediate states of these environments to be solvable by AI agents.
  • Write tests that verify agent solutions and balance strictness with correctness.
  • Iterate on tasks and tests based on QA feedback to ensure fairness and robustness.

Skills

Python
JavaScript
TypeScript
React
Docker
PostgreSQL
Kafka
Redis
Testing

Tools

Docker
PostgreSQL
Kafka
Redis
FastAPI
React

Job description

United States Digital Space LLC is building a dataset to evaluate AI coding agents by creating realistic developer environments with codebases, infrastructure, tickets, docs, and a believable development history. You will design tasks from intermediate states and craft what a solution must achieve to be considered solved.

The role emphasizes writing tests that accept all valid approaches while rejecting incorrect ones, and iterating on tasks based on QA feedback to keep evaluations fair and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer - AI Evaluation & Task Design
Senior Software Engineer - AI Evaluation & Task Design

Dorado • United States

Remote
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • United States

On-site
USD 41,000 - 69,000
AI Evaluation Engineer — Flexible Hours, High-Impact Testing
AI Evaluation Engineer — Flexible Hours, High-Impact Testing

Mindrift • United States

On-site
USD 41,000 - 69,000
AI Evaluation Engineer — Design & Validate Coding Tasks
AI Evaluation Engineer — Design & Validate Coding Tasks

Mindrift • Town of Texas (WI)

On-site
USD 55,000 - 69,000
Flexible schedule
Remote-friendly
AI Evaluation Architect - Project-Based Code Agent Testing
AI Evaluation Architect - Project-Based Code Agent Testing

Mindrift • New York (NY)

On-site
USD 41,000 - 69,000
Flexible schedule
Project-based work
Senior AI Evaluation Engineer – Agent Testing & Tasks
Senior AI Evaluation Engineer – Agent Testing & Tasks

Dorado • United States

Remote
AI Coding Task Designer & Evaluation Specialist
AI Coding Task Designer & Evaluation Specialist

Mindrift • New York (NY)

On-site
USD 34,440,000 - 48,216,000
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • Michigan

On-site
USD 41,000 - 69,000
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • New York (NY)

On-site
USD 41,000 - 69,000
Flexible schedule
Project-based work
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • Town of Texas (WI)

On-site
USD 55,000 - 69,000
Flexible schedule
Remote-friendly