A new study proposes an automated peer-review system using multiple large language models to evaluate AI-generated research papers across originality, rigor, clarity, and significance. Testing four leading AI Scientist frameworks against FARS benchmark papers revealed dramatic performance gaps, with FARS scoring more than twice as high as competitors on most evaluations, while establishing that multi-model LLM evaluation can reliably assess autonomous research quality.
Why it matters: As AI systems begin conducting autonomous research, developing reliable evaluation methods becomes critical for the industry to measure progress, compare competing systems, and ensure the quality of machine-generated scientific work.