Senior AI Agent Evaluation Engineer

Worky

United States

Remote

USD 55,000 - 69,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindrift is creating a dataset to evaluate AI coding agents by simulating real-world developer tasks. You will guide and evaluate the agent-written code, not write code from scratch, and help design tests that challenge Frontier models in coding scenarios.

This project-based role offers flexible scheduling and compensation up to $50/hr, depending on level and pace, with roughly 20 hours per task.

Qualifications

  • 5+ years in software development.
  • Strong stack knowledge: Python (FastAPI), JS/TS (React), Docker, Postgres, Kafka, Redis.
  • Experience writing tests (functional, integration).
  • English proficiency - B2+.

Responsibilities

  • Assist in building a dataset to evaluate AI coding agents.
  • Guide and evaluate agent-generated code rather than writing code from scratch.
  • Develop tests to validate model outputs and assess real-world developer tasks.

Skills

Software development experience
Testing experience
English proficiency

Tools

Python (FastAPI)
JavaScript/TypeScript (React)
Docker
Postgres
Kafka
Redis

Job description

Mindrift is creating a dataset to evaluate AI coding agents by simulating real-world developer tasks. You will guide and evaluate the agent-written code, not write code from scratch, and help design tests that challenge Frontier models in coding scenarios.

This project-based role offers flexible scheduling and compensation up to $50/hr, depending on level and pace, with roughly 20 hours per task.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Architect - Project-Based Code Agent Testing
AI Evaluation Architect - Project-Based Code Agent Testing

Mindrift • New York (NY)

On-site
USD 41,000 - 69,000
Flexible schedule
Project-based work
Senior AI Evaluation Engineer – Agent Testing & Tasks
Senior AI Evaluation Engineer – Agent Testing & Tasks

Dorado • United States

Remote
AI Agent Evaluation Engineer — Design & Test Coding Tasks
AI Agent Evaluation Engineer — Design & Test Coding Tasks

Mindrift • New York (NY)

Remote
USD 41,000 - 69,000
AI Evaluation Engineer — Design & Validate Coding Tasks
AI Evaluation Engineer — Design & Validate Coding Tasks

Mindrift • Town of Texas (WI)

On-site
USD 55,000 - 69,000
Flexible schedule
Remote-friendly
AI Agent Evaluation Engineer (Project-Based)
AI Agent Evaluation Engineer (Project-Based)

Mindrift • Germany (OH)

On-site
AI Task Architect & Evaluation Engineer
AI Task Architect & Evaluation Engineer

Mindrift • South Carolina

On-site
USD 41,000 - 69,000
Competitive compensation
Flexible scheduling
AI Evaluation Engineer — Flexible Hours, High-Impact Testing
AI Evaluation Engineer — Flexible Hours, High-Impact Testing

Mindrift • United States

On-site
USD 41,000 - 69,000
Python Engineer - AI Coding Agent Tests (Project)
Python Engineer - AI Coding Agent Tests (Project)

Dorado • United States

Remote
Software Engineer - AI Agent Evaluation
Software Engineer - AI Agent Evaluation

aitrainer • New York (NY), Northern (KY)

Hybrid
USD 57,000 - 81,000
AI Evaluation Engineer (Python, QA or Security)
AI Evaluation Engineer (Python, QA or Security)

Mindrift • Michigan

On-site
USD 41,000 - 69,000