Methodology

AIPR scores newly published preprints with a calibrated AI first read, then puts a human reviewer in charge of the result. Each paper is graded from 0 to 100 across novelty, rigor, applicability, clarity, and citation quality, with a written rationale for every score and a citation check against external databases. This page explains how that grading works and where the human stays in control.

Humans in the loop

Every AI-generated review on AIPR is a starting point for a human reviewer, not a final verdict. The model produces scores and per-dimension comments; a reviewer reads the draft, edits whatever needs editing, and decides what the author sees. The AI never sends output directly to an author. Reviewers are free to overrule any score the model assigned. What the author receives is the reviewer's final version, not the model's.

Borderline scores, citation concerns, and suspected misconduct are flagged for the reviewer's attention but are never decided by the AI alone.

A longer version of why this matters lives on the purpose page.

How It Works

Every week we look at thousands of newly published research papers. Each selected paper is read in full and graded across five scoring dimensions by the model, producing a structured draft for a human reviewer to take forward.

The grading pass produces per-dimension scores with a written rationale for each one, so the reviewer (and later the author) can see why the model assigned what it did. Reviewers can revise scores, rewrite rationales, or send the draft back for another pass before it is finalised.

Citations are checked against external academic databases to verify that referenced works exist and are attributed correctly. Within a given weekly cohort, papers are also compared against each other so that the relative ranking reflects differences the model can defend, not differences the model happened to imagine.

Once a reviewer approves the review, the paper and its full evaluation breakdown can be published to the leaderboard. Readers see the reviewer-approved version, not the raw model output.

The models

The weekly rankings are graded by OpenAI's GPT-5.4-mini. Full reviews of user-submitted manuscripts use GPT-5.4 as the reviewing model, with GPT-5.4-mini curating and assembling the output.

Scoring Dimensions

Each paper is scored on five dimensions, each from 0 to 100. Scores are combined into an overall score, then adjusted by model confidence.

Novelty

Evaluates the originality of the contribution. How new are the ideas, methods, or findings? Does the work introduce genuinely novel concepts or is it an incremental improvement?

Rigor

Assesses the methodological soundness and technical correctness. Are the proofs valid? Are experiments well-designed with proper controls? Are claims supported by evidence?

Applicability

Measures real-world relevance and potential impact. Can the methods be applied to practical problems? How broad is the potential audience? Does it solve a real need?

Clarity

Judges the quality of writing and presentation. Is the paper well-organized? Are the key ideas explained clearly? Are figures and tables informative?

Citation Quality

Evaluates the quality of references and related work coverage. Does the paper cite relevant prior work? Are comparisons fair and comprehensive? Are key baselines included?

How to read the scores

Scores are anchored so that a score around 70 corresponds to the acceptance bar at a strong venue. The distribution is deliberately concentrated in the upper range. A score below 70 is a signal to revise before submitting.

The score was validated against the public ICLR 2026 review record on OpenReview. The full study is on our publications page.

Frequently asked questions

Does AIPR replace human peer reviewers?

No. Every AI review is a draft for a human reviewer. The reviewer reads it, edits the scores and comments, and decides what the author sees. The model never sends its output to an author on its own.

Is it ethical to use AI in peer review?

AIPR is built for the case where a reviewer or editor stays accountable for the final review. The AI drafts a first read and flags concerns; the human makes the judgment. Confidential manuscript content is never used to train a model.

How does AIPR score a paper?

Each paper is read in full and scored from 0 to 100 on five dimensions: novelty, rigor, applicability, clarity, and citation quality. Every score comes with a written rationale, and the dimensions combine into an overall score adjusted for the model's confidence.

Are the references in a paper checked?

Yes. Citations are verified against external academic databases to confirm that referenced works exist and are attributed correctly.

Can I review my own paper before I submit it?

Yes. Authors can run a manuscript through AIPR to get a structured first read before sending it to a journal, then act on the feedback while there is still time to revise.

How are the weekly rankings produced?

Within each weekly cohort, papers are compared against one another so the ranking reflects differences the model can defend. A reviewer approves a review before the paper and its full breakdown are published to the leaderboard.

What does a score below 70 mean?

Scores are anchored so that around 70 corresponds to the acceptance bar at a strong venue, and the distribution is deliberately concentrated in the upper range. A score below 70 is a signal to revise the paper before submitting it.

Working on something with us?

Press, partnerships, bug reports, anything.

Contact us