Software Engineer – LLM Evaluation & Repository Validation

AI Trainer Jobs

United States

Remote

USD 69,000 - 138,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

AI Trainer Jobs is seeking software engineers to validate real-world repositories used to evaluate large language models. You will triage issues, set up reproducible environments, assess test quality, and check how well models fix real bugs.

You will review and triage issues in open-source GitHub projects, prepare repositories for evaluation with containerisation, and write up findings for researchers. This role is remote, independent, and deadline-driven.

Qualifications

  • 3+ years of software engineering experience.
  • Strong experience in at least one of Python, C++, C#, Go, Rust or Ruby.
  • Proficiency with Git, Docker and basic pipeline setup.
  • Ability to navigate complex code bases and run real-world projects locally.
  • Fluent written English.
  • Reliable, detail-oriented, and able to work independently and meet deadlines in a remote setting.

Responsibilities

  • Review and triage issues in open-source GitHub projects.
  • Prepare repositories for evaluation, including containerisation and environment automation.
  • Assess how thorough and well written the unit tests are.
  • Edit and run code locally to see how LLMs perform on bug-fixing tasks.
  • Help pick repositories and issues that are hard for LLMs.
  • Write up findings clearly for researchers and reviewers.
  • Follow project guidelines, take part in calibration and review sessions, and incorporate feedback to keep quality consistent.

Skills

Software engineering
English fluency

Tools

Git
Docker
CI/CD pipelines

Job description

Back

Software Engineer - LLM Evaluation & Repository Validation

Computer Science

Related Skills

No related skills for this job.

Available in

All Countries

About the Role

We are looking for software engineers to validate real-world repositories used to evaluate large language models. You will triage issues, set up reproducible environments, assess test quality, and check how well models fix real bugs.

Key Responsibilities
  • Review and triage issues in open-source GitHub projects.
  • Prepare repositories for evaluation, including containerisation and environment automation.
  • Assess how thorough and well written the unit tests are.
  • Edit and run code locally to see how LLMs perform on bug-fixing tasks.
  • Help pick repositories and issues that are hard for LLMs.
  • Write up findings clearly for researchers and reviewers.
  • Follow project guidelines, take part in calibration and review sessions, and incorporate feedback to keep quality consistent.
Who We’re Looking For
Required Qualifications
  • 3+ years of software engineering experience.
  • Strong experience in at least one of Python, C++, C#, Go, Rust or Ruby.
  • Proficiency with Git, Docker and basic pipeline setup.
  • Ability to navigate complex code bases and run real-world projects locally.
  • Fluent written English.
  • Reliable, detail-oriented, and able to work independently and meet deadlines in a remote setting.
Preferred Qualifications
  • Prior LLM research or evaluation projects.
  • Experience building or testing developer tools or automation agents.
  • Open-source contribution or evaluation experience.
Rights, Confidentiality and Data Handling

Submit only work that you created yourself or that you have the right to share. Do not submit material covered by an employer, client or third-party confidentiality agreement, or material you have no right to share. Do not submit export-controlled, military or defence-related technical data. Remove credentials, secrets and personal data from anything you submit. Rex may ask for evidence of ownership or licence, may verify submissions, and may reject or remove any submission.

Legal and Compliance Notice

All content you submit must be lawful, and you must have the right to share it. Do not submit confidential or proprietary material belonging to an employer, client or third party, or personal data without a lawful basis. Rex may verify identity, credentials, ownership and licensing; may reject or remove any submission; and may report unlawful submissions to the relevant authorities. Engagement and payment are subject to successful verification and to applicable sanctions, export-control, tax and employment laws. The contractor agreement will set out ownership and licence terms for the work you deliver (to be confirmed by Rex Legal).

Compensation

USD $50-100 per hour, adjusted by country of residence.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer
Software Engineer

AI Trainer Jobs • United States

Remote
USD 90,000 - 145,000
GitHub Contributor
GitHub Contributor

AI Trainer Jobs • United States

Remote
USD 69,000 - 138,000
Software Engineer (Ruby)
Software Engineer (Ruby)

turing • United States

Remote
USD 120,000 - 190,000
Fully remote
Cutting-edge AI projects
Remote LLM Evaluation & Repo Validation Engineer
Remote LLM Evaluation & Repo Validation Engineer

AI Trainer Jobs • United States

Remote
USD 69,000 - 138,000
JavaScript Engineer
JavaScript Engineer

turing • United States

Remote
USD 28,000 - 55,000
Fully remote environment
Cutting-edge AI projects
STEM Careers in the United States
STEM Careers in the United States

Rex.zone • United States

On-site
USD 41,328 - 68,880
C++ Systems Engineer – AI Model Evaluation & Code Review | Remote
C++ Systems Engineer – AI Model Evaluation & Code Review | Remote

Crossing Hurdles • United States

On-site
USD 82,656 - 137,760
Python Developer
Python Developer

turing • United States

Remote
USD 120,000 - 180,000
Work in a fully remote environment.
Opportunity to work on cutting-edge AI
Evaluations Engineer
Evaluations Engineer

Vibehackers • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 185,000
Relocation support
Health insurance
Lunch and dinner provided
+6
Embedded / Systems Engineer (C) – AI Code Analysis | Remote
Embedded / Systems Engineer (C) – AI Code Analysis | Remote

Crossing Hurdles • United States

On-site
USD 82,656 - 137,760