AI Research Peer Review Evaluator (ML/AI)

Lightly

Schweiz

Hybrid

EUR 83.000 - 138.000

Teilzeit

Vor 11 Tagen
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Fully remote
Part-time contractor role
Flexible hours
Work on cutting-edge AI evaluation

Zusammenfassung

Lightly AG, a Zurich-based AI company, is seeking researchers to support an AI evaluation project focused on scientific peer review. You will assess AI-generated reviews against expert human peer reviews of ML/AI research papers.

This is a remote, project-based contractor opportunity with flexible hours. You will evaluate reviews, identify gaps, and provide concise rationales using a structured rubric across multiple papers.

Qualifikationen

  • Advanced degree in ML/AI or related field; evidence of scholarship.
  • Ability to critically read ML/AI research papers including methodology and results.
  • Experience performing literature searches and verifying prior work dates.

Aufgaben

  • Read ML/AI papers to identify core contributions, methods, experiments and claims.
  • Review original human peer reviews to establish a baseline.
  • Evaluate AI-generated reviews against the baseline using a structured rubric.
  • Assess accuracy, depth, constructiveness, and novelty of AI reviews.
  • Identify hallucinations, unsupported claims, and missed issues.
  • Compare two AI reviews side-by-side and judge usefulness.
  • Search and verify relevant literature using Google Scholar, arXiv, etc.
  • Provide concise, evidence-based rationales per rubric.

Kenntnisse

Master's degree in ML/AI
PhD candidate or holder
Experience reviewing ML/AI papers
Literature search skills
Analytical writing
Familiar with NeurIPS ICML ICLR ACL CV
Attention to detail

Ausbildung

Master's degree in ML/AI/CS
PhD in ML/AI

Tools

Google Scholar
arXiv
Semantic Scholar

Jobbeschreibung

Lightly AG is a Zurich-based AI company and ETH/HSG spin-off, backed by Y Combinator and top-tier investors. Our machine learning and computer vision technology is trusted by global leaders in autonomous driving, medical imaging, and visual inspection.

We're looking for researchers with strong Machine Learning / AI backgrounds to support an AI evaluation project focused on scientific peer review. You'll evaluate reviews generated by agentic AI systems and compare them against expert human peer reviews of ML/AI research papers.

This is a remote, project-based contractor opportunity with flexible working hours.

Tasks

What you'll be doing

  • Read and scan ML/AI research papers to understand their core contributions, methodology, experiments, and claims
  • Review the original human peer reviews to establish an expert baseline for each paper
  • Evaluate AI-generated peer reviews against that baseline using a structured scoring rubric
  • Assess the technical accuracy, analytical depth, constructive value, and novelty/significance assessment of each AI review
  • Identify hallucinations, unsupported claims, missed technical issues, or valuable insights surfaced by the AI reviewers
  • Compare two AI-generated reviews side-by-side and determine where one provides stronger or more useful analysis
  • Search and verify relevant academic literature using sources such as Google Scholar, arXiv, or Semantic Scholar, including checking whether cited prior work was available before the paper's submission date
  • Provide concise, evidence-based rationales explaining your evaluation decisions and consistently apply the project rubric

The evaluation specifically looks at whether agentic AI reviewers can provide meaningful value beyond expert human reviewers - for example, by identifying relevant prior literature that humans missed, questioning important assumptions, or resolving inconsistencies using evidence.

Requirements

You're a strong candidate if you:

  • Have a Master's, PhD, or are currently pursuing graduate study in Machine Learning, Artificial Intelligence, Computer Science, Statistics, or a closely related technical field
  • Have contributed to at least one scientific/research paper, ideally as a first author, although co-authors and other substantial contributors are also welcome
  • Have experience critically reading ML/AI research papers, including evaluating methodology, experimental design, results, limitations, and scientific claims
  • Are familiar with major ML/AI research venues, such as NeurIPS, ICML, ICLR, ACL, CVPR, or comparable conferences and journals
  • Have prior academic peer-review experience, ideally for an ML/AI conference or journal - strongly preferred
  • Are comfortable conducting academic literature searches and verifying prior work, publication dates, citations, and novelty claims
  • Have strong analytical and written communication skills and can distinguish meaningful technical concerns from superficial criticism
  • Can provide clear, concise, evidence-based rationales for your decisions
  • Can consistently apply detailed evaluation guidelines and scoring rubrics across multiple papers and reviews
  • Have strong attention to detail, particularly when identifying factual inaccuracies or hallucinated technical claims
Benefits
  • Fully remote and flexible - work from anywhere
  • Part-time contractor role with flexible hours
  • Work directly on the evaluation of cutting-edge agentic AI systems for scientific research
  • Apply your ML/AI research expertise to help measure and improve the quality of AI-generated scientific peer review
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Remote ML Researcher: AI Peer-Review Evaluator
Remote ML Researcher: AI Peer-Review Evaluator

Lightly • Schweiz

Hybrid
EUR 83.000 - 138.000
Fully remote
Part-time contractor role
Flexible hours
+1
AI Science Writer, Nebius Academy (Contract)
AI Science Writer, Nebius Academy (Contract)

Lever, Inc. • Lavamünd

Vor Ort
EUR 83.000 - 106.000
Fixed monthly retainer
Professional growth opportunities
Flexible Europe-wide arrangements
SMB AI Power User — Competitive Evaluations
SMB AI Power User — Competitive Evaluations

Lever, Inc. • Österreich

Vor Ort
EUR 48.000 - 60.000
Remote work
Up to 40 hours per week
Senior Data Scientist — Evaluation & Scoring (Remote)
Senior Data Scientist — Evaluation & Scoring (Remote)

YO IT Consulting • Lavamünd

Vor Ort
EUR 81.900 - 145.600
Weekly payments via Stripe or Wise
Applied AI Engineer (all genders)
Applied AI Engineer (all genders)

Accenture DACH • Wien

Vor Ort
EUR 41.000 - 50.000
Attractive compensation package
Flexible working hours
Modern work environment
+2
Senior Applied AI Solutions Engineer
Senior Applied AI Solutions Engineer

Jobgether • Lavamünd

Vor Ort
EUR 174.000 - 304.000
Competitive base salary
Comprehensive benefits
Flexible work arrangements
+3
Senior Product Engineer (Darkplane, Agentic Platform)
Senior Product Engineer (Darkplane, Agentic Platform)

Lever, Inc. • Lavamünd

Vor Ort
EUR 127.000 - 201.000
Remote-first working environment
Office access Amsterdam
Office access New York
+2
Remote Document Analyst for AI Evaluation — Flexible Hours (Contract)
Remote Document Analyst for AI Evaluation — Flexible Hours (Contract)

JobsinAustria • Österreich

Remote
EUR 95.000 - 190.000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Stott and May • Wien

Hybrid
EUR 90.000 - 120.000
Hybrid work model
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • Wien

Vor Ort
EUR 90.000 - 125.000