paperchase.ai/evaluation rev. 2026-07

How We Evaluate Our AI

Researchers are right not to trust AI tools by default.

Here is how PaperChase earns it:

Stance assignments are graded against evaluation panels — sets of papers whose positions were judged by domain experts before the model saw them. Releases are gated: a change that regresses precision doesn't ship. Abstention is measured and expected; a model that never says "no clear stance" is guessing. Every placement in the product carries its reasoning, and every "Wrong camp?" click feeds the correction loop.

1.

Evaluation panels

Panels span multiple disciplines, and the experts grade first — the model's assignments are scored against theirs, not against its own confidence.

2.

Release gates

Every model change is scored against the panels before release. The default answer for a change is no; precision has to survive for it to ship.

3.

Abstention

Ambiguous papers land in a "No clear stance" panel, visible in the product — not quietly forced into the nearest camp.

4.

Reasoning and correction

The reasoning shown for each placement is the same evidence you'd use to check it yourself — quote-level, one click from every camp.

Camp drill-down showing a stance reason and its confidence
Fig. 1 — stance reasoning and confidence as shown in the product.
5.

Emerging axes

Emerging-axis detection ships with an explicit hedge, because taxonomy detection is younger than stance assignment and we say so. Where the system is less sure, the interface says less.

published metrics
stance precision
pending publication
abstention rate
pending publication
panel coverage
pending publication

Figures appear here when approved for release. Until then: graded against expert panels.

We publish this because our users referee papers for a living. The standard you hold the literature to is the standard you should hold your tools to.