Skip to main content
Back to timeline
arXivSource publication:

After rewriting 60 ICLR submissions into 1,260 versions, AI reviewers show 'false robustness'; SciCore's dual branch lifts ICC to 0.775

Synopsis

The authors formulate Rhetorical Robustness as a joint requirement (stable judgments under content-preserving rewrites plus discrimination across papers), build RobustReview from 60 anonymized ICLR 2026 submissions under 10 rhetorical conditions for 1,260 full manuscripts, and evaluate 30 reviewer configurations, finding that low rewrite sensitivity often coincides with score collapse across papers ('false robustness') and that human alignment ranks reviewers differently; they then introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted structured science core, achieving a leading joint stability-discrimination profile in the primary GPT-5.5 comparison (ICC 0.775, SPR 0.652, discriminability 0.

AI-generated editorial illustration: A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Interpretation

The paper formulates rhetorical robustness as two complementary requirements: within-paper stability across content-preserving rewrites and between-paper discrimination, with human alignment measured separately as a distinct dimension. Prior AI-review evaluations mainly asked about human agreement, review quality, or score prediction, treating rewrite sensitivity as an isolated bias; here stability and discrimination are bound into one joint metric family (MAD and Drift SD alongside ICC, SPR, and discriminability). The benchmark comprises 60 anonymized ICLR 2026 submissions, 10 sampled from each of six mean human-score intervals, with 10 rhetorical conditions each instantiated by GPT-5.5 and Claude Opus 4.8, yielding 1,260 full manuscripts; rewrites protect citations, figures, and compilation structure, and automated plus human fidelity audits indicate core technical content is largely preserved across five dimensions.

The benchmark exposes 'false robustness': very low rewrite-induced score drift can coincide with collapsed discrimination across papers. This shows that within-paper stability alone can mistake near-constant scoring for robustness, so robustness must be read jointly with discrimination. Gemini-3.5-Flash-Lite under Standard has the lowest MAD among existing configurations (0.205) but ICC of only 0.199 and discriminability 0.511, close to chance; DeepReviewer and OpenJudge show relatively low drift yet ICC of only 0.161 and 0.239; in the science-core branch, Core-Standard and Core-Strict assign constant confidence scores with zero SPR.

Human alignment and rhetorical robustness rank reviewer configurations differently, and a content-focused prompting protocol does not consistently improve robustness across backbones. Human agreement therefore cannot substitute for matched rewrite evaluation, and content-focused review is not solved by prompting alone. GPT-5.5 under Strict has the strongest human alignment (H-MAE 1.078, Spearman 0.529), whereas Persistent has weaker human alignment but higher ICC, SPR, and discriminability; Persistent improves all five robustness metrics only for GPT-5.5, while Claude Sonnet 5 lowers MAD and Drift SD but reduces ICC, SPR, and discriminability, and other backbones change mixedly.

SciCore averages a full-manuscript judgment with a science-core judgment, achieving a leading joint stability-discrimination profile in the primary GPT-5.5 comparison while maintaining competitive human alignment. Unlike instructions to disregard rhetoric, the method first transforms the input into a content-normalized structured scientific record while retaining the manuscript branch; ablations show direct science-core review beats reconstructing a paper first, and review policy still shifts the stability-discrimination balance when the representation is held fixed. SciCore leads on ICC 0.775, SPR 0.652, and discriminability 0.726, with the lowest H-MAE (1.072) and second-highest Spearman (0.488); relative to Manuscript-Strict, fusion reduces MAD from 0.766 to 0.476 and raises ICC from 0.632 to 0.775, with paired paper-family bootstrap support on all five robustness metrics; science-core representation diagnostics show matched similarity of 0.978 versus cross-paper controls of 0.616-0.619.

Perspective

This work is aimed at researchers and conference workflows that use or design AI review assistance: it provides a reusable controlled full-manuscript rewrite benchmark and a joint metric set so that 'does the judgment drift under rewriting' and 'can papers still be told apart' can be measured together; SciCore offers an implementation path that keeps full-manuscript review while adding a content-normalized judgment, and equal weighting needs no tuning and no extra fusion call, making it a natural default. The results apply to ICLR 2026-style full submissions with matched arXiv sources under the evaluated models and protocols; the cross-backbone experiments characterize the science-core branch rather than the fused reviewer, and the final fusion is evaluated with GPT-5.5.

Rewrites are designed to preserve reported scientific content, but the authors state exact equivalence cannot be guaranteed, the automated fidelity audit identifies a nonzero mismatch rate under complex conditions, and residual content differences may contribute to observed score variation; mean human scores provide only a limited external reference and do not capture disagreement among reviewers; science-core extraction can omit or misread details, and equal-weight fusion assumes the two branch scores are comparable on a shared scale; the cross-backbone experiments vary extraction and review together, do not isolate the two stages, and do not test the final fusion across backbones; the main experiments use single-review execution, so robustness metrics include stochastic generation variability. In addition, some appendix tables in the loaded text (such as the bootstrap interval tables and parts of the complete results tables) are empty, so those specific numbers cannot be summarized here.

Sources