Remote LLM Evaluation & Repo Validation Engineer

AI Trainer Jobs

United States

Remote

USD 69,000 - 138,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

AI Trainer Jobs is seeking software engineers to validate real-world repositories used to evaluate large language models. You will triage issues, set up reproducible environments, assess test quality, and check how well models fix real bugs.

You will review and triage issues in open-source GitHub projects, prepare repositories for evaluation with containerisation, and write up findings for researchers. This role is remote, independent, and deadline-driven.

Qualifications

  • 3+ years of software engineering experience.
  • Strong experience in at least one of Python, C++, C#, Go, Rust or Ruby.
  • Proficiency with Git, Docker and basic pipeline setup.
  • Ability to navigate complex code bases and run real-world projects locally.
  • Fluent written English.
  • Reliable, detail-oriented, and able to work independently and meet deadlines in a remote setting.

Responsibilities

  • Review and triage issues in open-source GitHub projects.
  • Prepare repositories for evaluation, including containerisation and environment automation.
  • Assess how thorough and well written the unit tests are.
  • Edit and run code locally to see how LLMs perform on bug-fixing tasks.
  • Help pick repositories and issues that are hard for LLMs.
  • Write up findings clearly for researchers and reviewers.
  • Follow project guidelines, take part in calibration and review sessions, and incorporate feedback to keep quality consistent.

Skills

Software engineering
English fluency

Tools

Git
Docker
CI/CD pipelines

Job description

AI Trainer Jobs is seeking software engineers to validate real-world repositories used to evaluate large language models. You will triage issues, set up reproducible environments, assess test quality, and check how well models fix real bugs.

You will review and triage issues in open-source GitHub projects, prepare repositories for evaluation with containerisation, and write up findings for researchers. This role is remote, independent, and deadline-driven.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
Remote Tech Lead Go Engineer for AI Code Evaluation
Remote Tech Lead Go Engineer for AI Code Evaluation

turing • United States

Remote
USD 60,000 - 140,000
Fully remote environment
Flexible hours
Remote Code Review Engineer — Evaluation Specialist
Remote Code Review Engineer — Evaluation Specialist

AI Trainer Jobs • United States

Remote
USD 158,000 - 200,000
Remote AI Code Evaluation Engineer
Remote AI Code Evaluation Engineer

AI Trainer Jobs • United States

Remote
USD 90,000 - 145,000
Remote AI/ML Code Reviewer & Debug Specialist
Remote AI/ML Code Reviewer & Debug Specialist

AI Trainer Jobs • United States

Remote
USD 165,000 - 193,000
Remote work
Remote ML Code Reviewer & Debugging Engineer
Remote ML Code Reviewer & Debugging Engineer

AI Trainer Jobs • United States

Remote
USD 165,000 - 193,000
Remote LLM Evaluation Specialist - Part-Time
Remote LLM Evaluation Specialist - Part-Time

OpenTrain AI, Inc. • Northern (KY)

Hybrid
USD 45,000 - 65,000
Software Engineer – LLM Evaluation & Repository Validation
Software Engineer – LLM Evaluation & Repository Validation

AI Trainer Jobs • United States

Remote
USD 69,000 - 138,000
Remote LLM Research Scientist – Pre/Post-Training Evaluator
Remote LLM Research Scientist – Pre/Post-Training Evaluator

AI Trainer Jobs • United States

Remote
USD 138,000 - 165,000
Remote Senior AI/ML Engineer – LLM Training & Evaluation
Remote Senior AI/ML Engineer – LLM Training & Evaluation

Prolific • Milwaukee (WI)

On-site
USD 100,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home