Georgetown launches an AI referee leaderboard for research papers
Georgetown University’s AI, Analytics, and the Future of Work Initiative has put a live leaderboard online that asks a frontier AI model to evaluate finance and economics working papers. The project, titled AI Referee Paper Leaderboard, presents the top 100 papers in several time windows and lets visitors open a detailed report for each entry.
The current system uses Anthropic’s Claude Opus 4.8 as an AI referee. According to the project’s explanation, the model reads extracted paper content and scores each manuscript from 0 to 100 against five dimensions: significance, originality, correctness, data and methodology, and exposition. A report can also include perceived strengths and weaknesses, suggested improvements, a confidence assessment and information about what the model could actually read.
That last part is central to the project’s design. The leaderboard is not presented as a replacement for journal editors, conference committees or human reviewers. Georgetown describes it as a real-time research experiment: papers enter the FEN network, receive an initial assessment and can later be compared with conference and publication outcomes. The goal is to test whether an AI score contains an early signal, not to declare that a paper is correct or important.
The initiative also places clear limits on interpretation. Scores are snapshots tied to a particular manuscript version, model, rubric and evaluation procedure. A high result does not prove a paper’s accuracy, originality, methodological quality or publication prospects. The model may miss information, favour polished presentation or produce inconsistent judgements. Human expertise, replication and conventional peer review remain necessary, especially when an assessment could affect a researcher’s reputation or career.
For AI users and makers, the experiment matters beyond academic publishing. It shows a practical pattern for long-document evaluation: apply a fixed multi-part rubric, return structured feedback and disclose uncertainty about the material that was reviewed. It also exposes the risk of treating a model-generated number as an objective verdict. Companies building AI evaluation pipelines can borrow the transparency, but should validate any score against the outcome they actually care about. Georgetown’s leaderboard is therefore best understood as both a public demonstration of AI-assisted review and a test of how carefully such systems must be governed.