How We Evaluate Our AI
Researchers are right not to trust AI tools by default.
Here is how PaperChase earns it:
Stance assignments are graded against evaluation panels — sets of papers whose positions were judged by domain experts before the model saw them. Releases are gated: a change that regresses precision doesn't ship. Abstention is measured and expected; a model that never says "no clear stance" is guessing. Every placement in the product carries its reasoning, and every "Wrong camp?" click feeds the correction loop.
1.Evaluation panels
Panels span multiple disciplines, and the experts grade first — the model's assignments are scored against theirs, not against its own confidence.
2.Release gates
Every model change is scored against the panels before release. The default answer for a change is no; precision has to survive for it to ship.
3.Abstention
Ambiguous papers land in a "No clear stance" panel, visible in the product — not quietly forced into the nearest camp.
4.Reasoning and correction
The reasoning shown for each placement is the same evidence you'd use to check it yourself — quote-level, one click from every camp.
5.Emerging axes
Emerging-axis detection ships with an explicit hedge, because taxonomy detection is younger than stance assignment and we say so. Where the system is less sure, the interface says less.
Figures appear here when approved for release. Until then: graded against expert panels.
We publish this because our users referee papers for a living. The standard you hold the literature to is the standard you should hold your tools to.