When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Synopsis
Starting from Llama 3.1 8B, the study first fine-tunes a reviewer on official ICLR reviews from 2018–2023, then trains four successor models on ICLR 2024 data with synthetic reviews generated by that reviewer mixed at 0%, 33%, 66%, and 100%, finding that higher synthetic exposure compresses rating distributions and monotonically reduces same-paper and corpus-level semantic diversity (about 11% and 5%), a pattern the authors call scientific-judgment collapse, and introduces TrustReviewer, which mitigates this tendency through training-time corpus curation and test-time paired activation steering.
Interpretation
It identifies and tests scientific-judgment collapse in recursive AI-reviewer training: introducing synthetic reviews compresses rating distributions, and both same-paper and corpus-level semantic diversity decline monotonically as synthetic exposure rises. Prior recursive-training work focused mainly on distributional tail loss or model collapse in generated data; this work targets conditional scientific judgment and provides controlled evidence in the peer-review setting. Controlled experiment: initialization and training configuration are held fixed while only the synthetic-review proportion varies across four variants (0%, 33%, 66%, 100%) sharing filtering and optimization settings; evaluated on 2,000 held-out papers, with reported reductions of about 11% in same-paper semantic distance and about 5% in corpus-level semantic spread, while mean ratings change non-monotonically, indicating diversity compression rather than a systematic shift toward leniency or harshness.
It introduces TrustReviewer, an open-source system that reduces low-quality and semantically degenerate supervision through unified filtering and corpus curation at training time, and further mitigates residual collapse tendencies through paired activation steering at inference time. It combines two existing lines of work, data curation and activation steering, for the review-generation task, estimating steering vectors from representational differences between official and model-generated reviews of the same papers, without further training or additional expert annotation. The curated corpus contains 112,743 paper–review examples totaling about 1.9 billion tokens; steering vectors are estimated on 5,000 calibration pairs, and hyperparameters are selected on 100 separate validation pairs by exact recommendation match, choosing the final decoder layer and strength 1.0.
On the held-out paper set, TrustReviewer achieves the highest exact recommendation match among evaluated models at 75.40%, the lowest mean absolute distance at 1.079, and rating entropy of 2.18, above Llama (1.53), Qwen (1.94), and OpenReviewer (2.10), narrowing the gap to the official reference of 2.38. Relative to its Llama initialization (33.93% exact match) and the external specialized reviewer OpenReviewer (73.10%), it offers comparable results on both recommendation agreement and judgment diversity. Three independent generation runs per paper at temperature 0.6, top_p 0.95, and up to 4,096 new tokens, without fixed decoding seeds; metrics are exact matching and mean absolute distance from the average official rating, with missing or unparseable ratings never assigned default values.
Test-time paired activation steering adds gains without parameter updates or additional expert annotation: exact match rises from 73.85% to 75.40% (+1.55 percentage points), rating entropy rises from 2.13 to 2.18, corpus-level semantic spread expands, and mean absolute distance remains essentially unchanged. This shows inference-time representation intervention can complement training-time curation, and that its effect is not a uniform improvement across every diversity metric, since same-paper semantic distance decreases slightly. The same TrustReviewer checkpoint is compared with steering on and off; the vector is added only at the last token position and model parameters remain unchanged.
Perspective
The work addresses conferences and journals that use LLMs to generate or assist in writing reviews, especially where review text enters public data and may be used to train later models; it provides a controlled single-step recursive experimental framework and reusable training-time curation and inference-time steering procedures for builders of review models who want to preserve judgment diversity while improving recommendation alignment.
The authors note the experiment covers one recursive step, one base-model family, and one review domain, so whether collapse compounds over multiple generations or holds across model scales and fields remains open; diversity metrics are embedding-based proxies and do not directly establish that reviews identify different valid scientific concerns; official reviews serve as an institutional reference rather than ground truth and may contain errors, disagreement, or unobserved AI assistance; exact recommendation match and mean absolute distance measure agreement with the official rating distribution, not the correctness of scientific judgment. In addition, some table and figure values appear as placeholders in the loaded text, so precise reproduction would require consulting the original figures and tables.
