Lightly AG is a Zurich-based AI company and ETH/HSG spin-off, backed by Y Combinator and top-tier investors. Our machine learning and computer vision technology is trusted by global leaders in autonomous driving, medical imaging, and visual inspection.
We're looking for researchers with strong Machine Learning / AI backgrounds to support an AI evaluation project focused on scientific peer review. You'll evaluate reviews generated by agentic AI systems and compare them against expert human peer reviews of ML/AI research papers.
This is a remote, project-based contractor opportunity with flexible working hours.
Tasks
What you'll be doing
- Read and scan ML/AI research papers to understand their core contributions, methodology, experiments, and claims
- Review the original human peer reviews to establish an expert baseline for each paper
- Evaluate AI-generated peer reviews against that baseline using a structured scoring rubric
- Assess the technical accuracy, analytical depth, constructive value, and novelty/significance assessment of each AI review
- Identify hallucinations, unsupported claims, missed technical issues, or valuable insights surfaced by the AI reviewers
- Compare two AI-generated reviews side-by-side and determine where one provides stronger or more useful analysis
- Search and verify relevant academic literature using sources such as Google Scholar, arXiv, or Semantic Scholar, including checking whether cited prior work was available before the paper's submission date
- Provide concise, evidence-based rationales explaining your evaluation decisions and consistently apply the project rubric
The evaluation specifically looks at whether agentic AI reviewers can provide meaningful value beyond expert human reviewers - for example, by identifying relevant prior literature that humans missed, questioning important assumptions, or resolving inconsistencies using evidence.
Requirements
You're a strong candidate if you:
- Have a Master's, PhD, or are currently pursuing graduate study in Machine Learning, Artificial Intelligence, Computer Science, Statistics, or a closely related technical field
- Have contributed to at least one scientific/research paper, ideally as a first author, although co-authors and other substantial contributors are also welcome
- Have experience critically reading ML/AI research papers, including evaluating methodology, experimental design, results, limitations, and scientific claims
- Are familiar with major ML/AI research venues, such as NeurIPS, ICML, ICLR, ACL, CVPR, or comparable conferences and journals
- Have prior academic peer-review experience, ideally for an ML/AI conference or journal - strongly preferred
- Are comfortable conducting academic literature searches and verifying prior work, publication dates, citations, and novelty claims
- Have strong analytical and written communication skills and can distinguish meaningful technical concerns from superficial criticism
- Can provide clear, concise, evidence-based rationales for your decisions
- Can consistently apply detailed evaluation guidelines and scoring rubrics across multiple papers and reviews
- Have strong attention to detail, particularly when identifying factual inaccuracies or hallucinated technical claims
Benefits
- Fully remote and flexible - work from anywhere
- Part-time contractor role with flexible hours
- Work directly on the evaluation of cutting-edge agentic AI systems for scientific research
- Apply your ML/AI research expertise to help measure and improve the quality of AI-generated scientific peer review