Skip to main content
Back to timeline
arXivSource publication:

SAGO re-evaluates eleven LLMs by per-instance stability: no model generalizes uniformly, and cross-dataset variation can reverse rankings

Synopsis

The work introduces the Stability-Aware Generalization Objective (SAGO), which defines LLM generalization as per-instance behavioral stability of the same input under semantically equivalent variations, and evaluates eleven open- and closed-source models across six datasets on four behavioral axes—generation consistency, internal activations, confidence, and response mirroring—finding that every model exceeds a significance threshold on at least one axis, that the axes capture independent failure modes, and that cross-dataset variation can reverse model rankings.

AI-generated editorial illustration: Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

Interpretation

SAGO recasts generalization from aggregate accuracy to per-instance behavioral stability: for each prompt it builds a set of semantically equivalent variations, compares baseline and variant behavior along multiple behavioral axes, and aggregates normalized per-prompt RMS instability into a Stability Generalization Score (SGS). Prior work typically evaluates robustness within a single domain, a single rewrite family, or an aggregate benchmark score; this work places multiple datasets, multiple variation families, and multiple behavioral axes inside one per-instance analysis, and explicitly targets variability rather than a score that narrow training can improve. The paper gives a formal objective (Eqs. 1–2), normalization and RMS aggregation definitions, and a one-sided t-test against a non-zero tolerance threshold, with primary results at 5% and robustness reported at 1%, 5%, and 10%.

Four behavioral axes are instantiated: activation geometry, generation consistency, confidence and uncertainty, and response mirroring, spanning the path from internal representation to surface output. Hidden-state geometry, BLEU/BERTScore, sequence log-probability and token entropy, and stylistic accommodation were each developed in separate literatures; this work combines them within a single per-instance analysis and measures mirroring as a behavioral sensitivity without presupposing intent. Each axis has a concrete metric and scale factor, e.g. the activation axis uses cosine similarity of the last hidden state, generation consistency uses both BLEU and BERTScore-F1 with a RoBERTa-large backbone, and confidence retains both a length-sensitive and a length-invariant view; mirroring detectors are family-specific, with social-register mirroring judged by GPT-4o-mini and 10% of labels manually verified.

The evaluation shows generalization instability is widespread and heterogeneous: no model generalizes uniformly, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings. This directly challenges treating a single robustness score or a single behavioral metric as a proxy for generalization, and supplies concrete counterexamples: within the Gemma family G4-E4B is less generation-stable than its predecessors, while Q3.5-9B has the best generation consistency in the open-source pool yet the highest confidence and mirroring instability. Eleven models (eight open-source, three closed-source) across six datasets; -BLEU and -MR are significant for every model; -Ent stays within the acceptable threshold for all models (maximum about 0.012), indicating token-level distributional sharpness is largely decoupled from output-level stability; L-8B's -Cos and -MR vary across the six datasets enough to reverse rankings.

Neither scale nor closed-source status removes instability; it shifts instability to different behavioral dimensions. The paper reports that L-8B, G-7B, and Q-7B are equally or more unstable than their smaller siblings on most axes, ruling out the hypothesis that the failure mode resolves with parameter count; closed-source models, observable only on some axes due to API limits, show a different instability profile from open-source ones. The scale comparison rests on SGS contrasts among same-family models of different sizes; on the closed-source side only GPT-5.4 exposes the per-token probabilities needed for -Ent, and the activation and sequence log-probability axes are unavailable, so closed-source conclusions are confined to the observable axes.

Perspective

The framework applies where one needs to know whether a model's behavior stays stable across different expressions of the same goal and context, for example pre-deployment robustness profiling, model selection, and acceptance testing of training interventions; it yields behavioral evidence about where and how much behavior changes, not causal mechanisms. For practitioners, the most direct use is to read SGS as a profile per axis, per dataset, and per variation family rather than as one comparable score; the paper also recommends profile-based rather than rank-based comparison, since one model may be preferable when content stability matters most and another when confidence stability or reduced surface mirroring does.

The boundaries the paper itself states are worth watching: SAGO measures where and how much behavior changes but not causal mechanisms, leaving representational drift, decoding sensitivity, and mirroring entangled; datasets are English-only, so whether the same axis dissociations hold across languages, scripts, and culturally specific prompt conventions is untested; -BLEU captures lexical deviation and may reflect harmless paraphrasing, hence -BERT as a meaning-level complement; the 5% operating point is validated for typical per-prompt instability, while rare or worst-case failures may require larger samples; structural variants and social-register mirroring rely on GPT-4o-mini and may introduce stylistic bias, and limited API access confines closed-source conclusions to the observable axes. In addition, in the text read here some equations, figures, and appendix tables appear as placeholders or truncated spans, so the specific numbers in the sample-size ablation and some figure captions could not be fully checked, and details of the recommended sample size and the placement and strength ablations should be confirmed against the original.

Sources