Senior AI Code Evaluator (Python/TypeScript)

Turing

United States

Remote

USD 83,000 - 152,000

Part time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Turing is seeking experienced software engineers to evaluate and improve AI coding models. You’ll review AI-generated code and agent behavior across real-world repositories, judge correctness, and identify failure modes to help refine model performance.

In this contractor role, you’ll collaborate with researchers and engineers, design rubrics, and produce clear evaluation data and actionable recommendations to advance coding models.

Qualifications

  • 5+ years hands-on software engineering experience
  • Proficiency in Python, TypeScript/JavaScript, Go or another major production language
  • Experience reviewing AI-generated code and coding agents

Responsibilities

  • Evaluate AI-generated code across real-world repositories
  • Review agent behavior, tool usage, and code changes for correctness
  • Identify technical errors, weak approaches, and recurring model failure modes
  • Compare model outputs and explain why one solution is better than another
  • Create and refine rubrics and evaluation criteria for coding tasks
  • Produce high-quality evaluation data and recommendations for the team
  • Build and maintain pipelines and infrastructure for data generation, collection, and evaluation
  • Synthesize findings into clear write-ups and updates for the team
  • Collaborate with researchers and engineers to translate qualitative judgment into scalable processes
  • Share clear, actionable findings with AI researchers and engineers

Skills

5+ years hands-on software engineering
Python
TypeScript/JavaScript
Go
Code review
Written communication
AI coding tools
LLM evaluation

Tools

Git
CI/CD
LLM tooling

Job description

Turing is seeking experienced software engineers to evaluate and improve AI coding models. You’ll review AI-generated code and agent behavior across real-world repositories, judge correctness, and identify failure modes to help refine model performance.

In this contractor role, you’ll collaborate with researchers and engineers, design rubrics, and produce clear evaluation data and actionable recommendations to advance coding models.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Code Reviewer (Python/TypeScript)
Senior AI Code Reviewer (Python/TypeScript)

Turing • Seattle (WA)

On-site
USD 110,000 - 193,000
Remote AI Software Engineer - Code & Data Evaluation
Remote AI Software Engineer - Code & Data Evaluation

Turing • United States

Remote
USD 83,000 - 138,000
Software Engineer - Python/Typescript
Software Engineer - Python/Typescript

Turing • Seattle (WA)

On-site
USD 110,000 - 193,000
Java AI Code Evaluator (Contract, 10–40 hrs/wk)
Java AI Code Evaluator (Contract, 10–40 hrs/wk)

Turing • New York (NY)

On-site
USD 83,000 - 165,000
Remote AI Software Engineer: Code Quality & Evaluation
Remote AI Software Engineer: Code Quality & Evaluation

Turing • United States

Remote
USD 55,000 - 110,000
Remote Software Engineer - Python/Typescript
Remote Software Engineer - Python/Typescript

Turing • United States

Remote
USD 83,000 - 152,000
Remote AI Code Engineer — Benchmark & Validate Models
Remote AI Code Engineer — Benchmark & Validate Models

Turing • United States

Remote
USD 69,000 - 138,000
Remote AI Software Engineer - Data & Model Training
Remote AI Software Engineer - Data & Model Training

turing • San Francisco (CA)

Remote
USD 83,000 - 124,000
Senior AI Interaction Evaluator for Coding Agents
Senior AI Interaction Evaluator for Coding Agents

Precision Labs • Northern (KY)

Hybrid
USD 138,000 - 276,000
Remote C/C++ AI Software Engineer (Contractor)
Remote C/C++ AI Software Engineer (Contractor)

Turing • United States

Remote
USD 55,000 - 96,000
Flexible hours
1 month contract