Remote Part-Time Java Engineer — LLM Benchmarking

Jointaro

United States

On-site

USD 73,873 - 105,214

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jointaro is seeking a motivated Java Software Engineer for part-time, remote opportunities to evaluate Large Language Models (LLMs). You will contribute to high-impact research by identifying gaps and explaining model failures while developing coding benchmarks reflecting real-world scenarios.

This role requires a commitment of 15+ hours per week and offers a competitive pay starting at $65/hour, plus bonuses for task completion. Candidates must be located in the US with proper work authorization.

Qualifications

  • Experience with software development using Java.
  • Strong understanding of unit and integration testing practices.
  • Ability to provide structured feedback and optimize code.

Responsibilities

  • Develop and validate coding benchmarks from real-world repositories.
  • Ensure comprehensive unit and integration tests for verification.
  • Maintain scalability of benchmark task distribution.
  • Provide feedback on solution quality and optimize benchmark code.
  • Document processes for reproducibility.

Skills

Java
Software Engineering
Unit Testing
Integration Testing

Job description

Jointaro is seeking a motivated Java Software Engineer for part-time, remote opportunities to evaluate Large Language Models (LLMs). You will contribute to high-impact research by identifying gaps and explaining model failures while developing coding benchmarks reflecting real-world scenarios.

This role requires a commitment of 15+ hours per week and offers a competitive pay starting at $65/hour, plus bonuses for task completion. Candidates must be located in the US with proper work authorization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Part-Time JS Engineer — AI Benchmarking & LLM QA
Remote Part-Time JS Engineer — AI Benchmarking & LLM QA

Jointaro • United States

On-site
USD 74,000 - 105,000
LLM Evaluations Engineer — Remote Benchmarking & Tools
LLM Evaluations Engineer — Remote Benchmarking & Tools

Jaide Health • United States

On-site
USD 90,000 - 130,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance for you and dependents
+4
Remote Part-Time Python AI Engineer — LLM Benchmarking
Remote Part-Time Python AI Engineer — LLM Benchmarking

Jointaro • United States

On-site
USD 90,000 - 124,000
Remote AI Math Evaluator & LLM Benchmark Designer
Remote AI Math Evaluator & LLM Benchmark Designer

United States Digital Space LLC • United States

Remote
Fully remote
AI projects
Contract extension potential
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000
AI Engineer - LLM Training & Evaluation (Remote)
AI Engineer - LLM Training & Evaluation (Remote)

Prolific • Memphis (TN)

Hybrid
USD 90,000 - 130,000
Competitive pay rates
Flexible hours
Ability to work from home
JavaScript Software Engineer
JavaScript Software Engineer

Jointaro • United States

Remote
USD 74,000 - 105,000
ML Engineer: LLM Evaluation & Observability
ML Engineer: LLM Evaluation & Observability

Gleanwork • Mountain View (CA)

Hybrid
USD 200,000 - 300,000
Health insurance
401(k) plan
Home office improvement stipend
+3
Applied Research Engineer – LLMs & Code Gen (Remote)
Applied Research Engineer – LLMs & Code Gen (Remote)

Jaide Health • United States

On-site
USD 120,000 - 150,000
Fully remote work & flexible hours
37 days/year of vacation & holidays
Health insurance allowance
+3
LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2