Code Quality Engineer for LLM Evaluation

OpenTrain AI, Inc.

United States

Remote

USD 34,000 - 69,000

Part time

47 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

OpenTrain AI, Inc. seeks a part-time contractor to review AI-generated code used to train and evaluate large language models, across languages and software scenarios.

You will identify defects, help create benchmarks, and offer feedback to improve reference solutions. This role is remote and requires strong written English, with 20+ hours weekly.

Experience across languages, debugging, and cross-language architecture is highly valued for this position.

Qualifications

  • Proven ability to review and critique AI-generated code for correctness and security.
  • Experience with production software development and modern engineering practices.
  • Ability to evaluate multiple language implementations and architectures.

Responsibilities

  • Review AI-generated code and software changes for correctness, security, reliability, scalability, readability, and maintainability.
  • Debug complex codebases and identify root causes of technical problems.
  • Compare implementations and select or create solutions that meet technical requirements.
  • Write clear technical feedback and contribute to evaluation criteria and benchmarks.

Skills

Python
JavaScript
TypeScript
C++
Go
Git

Education

Seven years of professional software engineering experience

Tools

Git
Code reviews
APIs
Databases

Job description

The Work

You will review AI-generated code used to train and evaluate large language models. Your work will cover different programming languages and software situations, including bug fixes, new features, refactoring, API integrations, configuration changes, and database operations.

You will assess whether code works as intended and identify defects, logic errors, missing pieces, edge-case failures, performance problems, security risks, and architecture weaknesses. You will compare possible implementations, improve code into reliable reference solutions, explain your recommendations, and help create technical rubrics and coding benchmarks.

  • Review AI-generated code and software changes for correctness, security, reliability, scalability, readability, and maintainability.
  • Debug complex codebases and identify the root causes of technical problems.
  • Compare implementations and select or create the solution that best meets the technical requirements.
  • Write clear technical feedback and contribute to evaluation criteria, coding benchmarks, and improved LLM evaluation methods.
What It Pays And Takes

The role details do not list a pay rate. This is a remote, part-time contract role for an individual contributor supporting AI training and engineering evaluation work.

  • Pay: Not provided in the role details.
  • Time: 20+ hours per week.
  • Location: Worldwide and remote.
  • Language: Written English proficiency is required.
  • Experience: The listing is marked entry level in the source fields, while the role description requires at least seven years of professional software engineering experience.
  • Programming: Strong proficiency in at least one of Python, JavaScript, TypeScript, Java, C++, Go, C#, Ruby, PHP, or Rust.
  • Technical knowledge: Production software development, debugging, code review, clean code, modular architecture, abstraction, error handling, data structures, algorithms, APIs, databases, and application architecture.
  • Practices: Experience with collaborative code reviews, Git, and modern software engineering methods.
  • Helpful background: Experience evaluating AI-generated code, creating technical rubrics, contributing to software engineering benchmarks, or working across multiple languages and architecture patterns.
About AI Training Work

AI training is the human work behind systems that generate and understand code, text, images, and other data. People review examples, rate model outputs, and provide clear corrections so AI systems become more useful and reliable; OpenTrain helps people find and build careers in this field.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Open Source Code Review Maintainer
Open Source Code Review Maintainer

OpenTrain AI, Inc. • Northern (KY)

Hybrid
USD 207,000 - 413,000
Software Engineer AI Code Evaluation
Software Engineer AI Code Evaluation

OpenTrain AI, Inc. • Northern (KY)

Hybrid
USD 138,000 - 207,000
Remote contractor
Part-time engagement
Hourly pay $100–$150 USD
+1
LLM Evaluation Specialist
LLM Evaluation Specialist

OpenTrain AI, Inc. • Northern (KY)

Hybrid
USD 45,000 - 65,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

On-site
USD 68,880 - 206,640
Remote Code Quality Engineer for AI Evaluation
Remote Code Quality Engineer for AI Evaluation

OpenTrain AI, Inc. • United States

Remote
USD 34,000 - 69,000
Legal AI Response Evaluation Lawyer
Legal AI Response Evaluation Lawyer

OpenTrain AI, Inc. • Northern (KY)

Hybrid
USD 124,000 - 193,000
AI Red Team Engineer for LLMs
AI Red Team Engineer for LLMs

OpenTrain AI, Inc. • United States

Remote
USD 47,000 - 63,000
Remote work
Worldwide eligibility
Legal AI Evaluation Expert
Legal AI Evaluation Expert

OpenTrain AI, Inc. • United States

Remote
USD 2,066,000 - 3,444,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • San Francisco (CA)

On-site
USD 68,880 - 206,640
Machine Learning Engineering Evaluator
Machine Learning Engineering Evaluator

OpenTrain AI, Inc. • United States

Remote
USD 138,000 - 207,000
Remote work
Flexible schedule
Portfolio building