AI Evaluator - Python (Freelance Opportunity)

Biz Tech Consultants

United States

Remote

USD 55,000 - 110,000

Part time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Biz Tech Consultants is seeking experienced software engineers to work on RLHF projects that support the training and evaluation of AI coding agents. You will design realistic technical tasks and review AI-generated responses for correctness and clarity, providing evidence-based feedback.

Ideal candidates have 4–7 years building production systems, strong JavaScript skills, and hands-on experience with Docker, databases, and remote or international teams.

Qualifications

  • 4 to 7 years of experience building production systems.
  • Strong hands-on skill in Javascript.
  • Comfortable with Docker and command line tools.
  • Experience with one relational and one NoSQL database.
  • Evidence of a clear before/after metric from past work.
  • Experience working with remote or international teams.
  • Exposure to secure data handling or compliance work.

Responsibilities

  • Design technical tasks across software development, debugging, data processing, security and ML; set up environments and tests as needed.
  • Review AI-generated responses for correctness, code quality and adherence to instructions.
  • Provide clear, evidence-based written feedback and comparative ratings across responses.
  • Validate task or evaluation difficulty against AI model outputs and revise as needed.

Skills

Technical writing
Code review
Task authoring
Remote collaboration

Tools

JavaScript
Node.js
Python
Go
Rust
Docker
Git
pytest
LLM API
Kubernetes
Kafka
Redis
gRPC
CI/CD
System Design

Job description


Job Description We are looking for experienced software engineers to work on RLHF projects that support the training and evaluation of AI coding agents. Depending on the project, the work can involve designing realistic technical tasks, or reviewing and rating AI generated responses for correctness, clarity, and reliability. It suits engineers who enjoy precise technical writing and careful quality review as much as building.

Key Responsibilities
  1. Design technical tasks across areas such as software development, debugging, data processing, security, and machine learning, and set up any supporting environment and automated tests needed to validate them, where the project calls for task authoring.
  2. Review AI generated responses to technical tasks, assessing correctness, code quality, and adherence to instructions.
  3. Provide clear, evidence based written feedback and comparative ratings across responses.
  4. Validate task or evaluation difficulty against AI model outputs and revise based on review feedback.

Desired Candidate ProfileHighly Required
  1. 4 to 7 years of experience building production systems, not just internal tools.
  2. Strong hands-on skill in Javascript
  3. Comfortable working with Docker and command line tools.
  4. Comfortable working with one relational and one NoSQL database.
  5. Can point to a clear before and after metric from past work.
  6. Experience working with remote or international clients or teams.
  7. Some exposure to secure data handling or compliance work.
Preferred
  1. Has integrated an AI or LLM API into a real, shipped feature.
  2. Experience in AI training and evaluation work.
  3. Prior experience debugging, reviewing code, and writing documentation or test cases.

Employment Type Contractual/FreelanceKey Skills Python, Java, JavaScript, Node.js, Go, Rust, Docker, Git, pytest, AI Integration, LLM, NoSQL, CI/CD, System DesignNice to Have Kubernetes, Kafka, RabbitMQ, gRPC, Redis

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluator - Freelance Opportunity (Python)
AI Evaluator - Freelance Opportunity (Python)

Biz Tech Consultants • United States

Remote
USD 120,000 - 180,000
AI Evaluator - JavaScript (Freelance Opportunity)
AI Evaluator - JavaScript (Freelance Opportunity)

Biz Tech Consultants • United States

Remote
USD 83,000 - 152,000
AI Evaluator - Typescript (Freelance Opportunity)
AI Evaluator - Typescript (Freelance Opportunity)

Biz Tech Consultants • United States

Remote
USD 83,000 - 165,000
RLHF AI Evaluator & Task Designer (Python)
RLHF AI Evaluator & Task Designer (Python)

Biz Tech Consultants • United States

Remote
USD 55,000 - 110,000
AI Code Evaluator (JavaScript) - RLHF Tasks, Contract
AI Code Evaluator (JavaScript) - RLHF Tasks, Contract

Biz Tech Consultants • United States

Remote
USD 83,000 - 152,000
AI Evaluation & Task Design Engineer
AI Evaluation & Task Design Engineer

Biz Tech Consultants • United States

Remote
USD 120,000 - 180,000
AI Evaluation Specialist
AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Austin (TX)

On-site
USD 68,191 - 97,120
Competitive pay
Flexible schedule
Experience in advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Alabama

On-site
USD 90,921 - 129,494
Flexible working hours
Competitive pay up to $80/hour
Experience in advanced AI projects