Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation
Synopsis
This work introduces information satisfaction—a query paired with a reader persona—as an axis of summarization evaluation, and through five perturbation tests plus an expert human evaluation finds that traditional metrics such as ROUGE and BERTScore and LLM-as-judge metrics such as Llama-3.3-70B and Prometheus-7B mostly fail basic perturbation checks and agree with reader preferences at near-chance levels, indicating that existing metrics are insufficient measures of how well a summary serves a specific reader's informational needs.
Figure 1: Annotation workflow example. Steps 1c-1d continue for each query assigned to the annotator.
· Page 5Interpretation
Introduces information satisfaction as a reader-centered evaluation axis, operationalized as a query paired with a persona describing the reader's role, domain, and information needs, to capture whether a summary resolves the specific informational need that motivated the reader. Prior evaluation treats quality as a fixed property of the summary or focuses on stylistic reader preferences; this work explicitly incorporates reader background as a comparatively stable, reusable signal across queries. A conceptual contribution operationalized through perturbation experiments on 4 scientific summarization corpora (arXiv, PubMed, SciTLDR, eLife, each subsampled to N=50 documents) and a human evaluation with 20 annotators.
Most popular metrics fail basic perturbation tests: on the different-audience rewrite, nearly every metric is insensitive or even anticorrelated; on incremental addition, ROUGE-1/2/L, BLEU, chrF, METEOR, and compression show negative correlations between −0.56 and −0.65, meaning partial summaries score higher than complete ones. Extends the previously documented length-sensitivity failure mode of ROUGE to BERTScore and most LLM-judge dimensions, and systematically tests the audience-shift dimension for the first time. Based on Spearman rank correlation and monotonicity statistics across 4 corpora with N=50 documents, with explicit pass thresholds (directional |ρ|≥0.4, stability |ρ|≤0.2, monotonicity M≥0.5).
LLM-as-judge metrics are highly inconsistent across judge models: dimensions that pass under Llama-3.3-70B routinely fail or invert under Prometheus-7B; on the different-audience test, Prometheus-7B's fluency, informativeness, overall, and FActScore even show strong positive correlations of 0.80–1.00. Reveals that LLM-as-judge conclusions are highly sensitive to the choice of judge model rather than being a stable quality signal. The same metric set is run under two judge models in parallel, directly comparing their Spearman ρ and match-rate M.
The expert human evaluation shows that no metric family exceeds chance agreement with reader preferences: traditional metrics range from α=−0.029 to 0.173, Prometheus-based judgments from −0.150 to 0.158, and Llama-based judgments from −0.179 to 0.053; the persona precision and persona recall variants designed specifically to capture information satisfaction perform among the worst (α=−0.150 for both under Prometheus). Moves metric evaluation from correlation with generic quality annotations to correlation with a specific reader's preferences under a specific query and persona, and shows that adding persona information alone is not sufficient. 140 completed query sessions from 20 annotators yielding 420 pairwise comparisons, using Krippendorff's α to measure agreement between metrics and annotator choices.
Perspective
This work targets scientific document summarization, evaluating information satisfaction under a query plus reader persona, and applies to research and engineering practice that generates or evaluates summaries for readers with different backgrounds (e.g., researchers, clinicians, students). The perturbation testing framework, annotation platform, and human-annotated data are released, enabling follow-up work to design new evaluation metrics or extend to other domains and languages.
In the human evaluation, annotators were recruited through professional networks with a sample of 20, so whether their preference distribution represents broader reader populations remains an open question; the perturbation tests use Llama-3.3-70B to generate rewrites, and rewrite quality itself may affect metric responses; the failure of persona precision and persona recall suggests the bottleneck may lie in how satisfaction is measured (e.g., weighting information by importance, modeling what the reader already knows, capturing comparative judgments between summaries) rather than in whether persona information is provided, and these directions await further research.
