Preprints.ai
<- All evidence pages
Model screening panel

Calibration corpus & review-12b

9,279 reviews assembled from eLife, preprints.ai self-generated, and PREreview, used to train a calibration model that maps the 11-agent panel output to a final grade. The model is in training and not yet wired into the production pipeline.

9,279 reviews CC-BY 4.0 (PREreview slice) review-12b: in training
eLife reviews
6,411
peer-reviewed, published
preprints.ai model outputs
1,570
self-generated, used for consistency only
PREreview
1,298
harvested from Zenodo, CC-BY 4.0

Pipeline today

The shipping production pipeline is described on the methodology page. Eleven specialist roles review routed manuscript context; a calibrated judge produces the Evidence letter subject to deterministic hard gates, while Trust is rule-based and Novelty is corpus-binned with literature grounding when available.

An earlier deterministic label lookup was built from roughly 800 historical eLife reviews. That mapping is useful historical calibration material, but it is not a gold-standard validation of the current three-axis production score.

What review-12b is

review-12b is an in-training calibration model whose only job will be to map the structured output of the 11-agent panel to the same grade scale, learned end-to-end against the 9,279-review corpus rather than via the hand-tuned lookup. It is a shape-matching layer, not a replacement reviewer.

Status

review-12b is not integrated into the live grade pipeline. Current Evidence grades use the calibrated judge and hard gates described on the methodology page. If review-12b is evaluated for production, we will publish agreement, calibration error, grade-shift histograms, and a human-audited disagreement set before rollout.

Why we have not shipped it yet. A calibration model that disagrees with the deterministic lookup must justify itself before it ships. The disagreement set has not yet been audited end-to-end by a human reviewer. We would rather miss a launch date than ship a model whose disagreements we cannot explain.

Caveats — what this doesn't measure

Code & attribution

Pipeline orchestration: agents/agentic_review.py. PREreview reviews are reproduced under CC-BY 4.0; eLife public reviews are reproduced under their open licensing.