Wharton Study Links Self-Ranked AI Papers to Higher Citation Counts

Researchers at the Wharton School found that authors' own rankings of their AI papers correlate strongly with future academic impact, outperforming traditional peer review metrics.
Key points
- Wharton researchers found that papers ranked highest by their own authors received twice as many citations as those ranked lowest.
- The study analyzed 1,527 unique papers from the 2023 International Conference on Machine Learning to test the validity of self-assessment.
- A new system flags papers with large gaps between self-ranks and review scores for additional scrutiny by area chairs.
A new study from the Wharton School suggests that asking authors to rank their own artificial intelligence papers can help identify high-impact research that traditional peer review might miss. The findings, reported by The Daily Pennsylvanian, challenge the assumption that external reviewers are the sole arbiters of scientific quality in a rapidly growing field.
The research was prompted by a widening gap between the volume of AI submissions and the limited pool of experienced peer reviewers. By analyzing data from the 2023 International Conference on Machine Learning, the team explored whether self-assessment could serve as a reliable signal of a paper's potential influence, offering a new layer to the evaluation process.
Self-Rankings Predict Citation Success
In the experiment, authors with multiple submissions were asked to rank their papers by perceived scientific quality before reviews were released. This constraint forced objectivity, as authors could not assign a perfect score to every submission. The results showed that papers ranked highest by their creators received an average of twice as many citations over the following 16 months compared to those ranked lowest.
This correlation held true for both accepted and rejected papers. Among the most cited studies in the sample, 77% were ranked first by at least one of their authors. These self-rankings predicted future citation counts more accurately than the initial peer-review scores, indicating that authors may have a sharper internal gauge of their work's significance.
Flagging Papers for Closer Scrutiny
The researchers do not propose replacing peer review with self-ranking. Instead, they suggest using the discrepancy between an author's self-score and the reviewers' score to flag papers for additional attention. Area chairs would receive an alert only when the gap between these scores is significant, prompting a deeper examination of the work.
This system is designed to reduce manipulation incentives. If an author ranks a weaker paper highly, it triggers a discrepancy that draws scrutiny rather than guaranteeing acceptance. Area chairs can then decide to recruit additional reviewers or engage in further discussion with the authors, ensuring that potentially overlooked high-quality work is not dismissed by a single review cycle.
Implementation and Trade-Offs
The International Conference on Machine Learning has incorporated this approach into its 2026 review process. A randomized experiment found that area chairs who saw these discrepancy categories wrote 71% more comment text per paper and communicated more frequently with reviewers. This suggests the method encourages more thorough engagement with the material.
However, citation counts are not a perfect proxy for quality. Factors such as topic popularity, release timing, and author visibility can influence how often a paper is cited. The study acknowledges that while self-rankings are a strong predictor of impact, they should be viewed as one tool among many in the complex landscape of academic evaluation.






