LUMOS traces parametric knowledge with OLMo 2's transparent corpus: rare facts are 84% separable internally but only 54% expressed behaviorally, and self-reflection falls to 49% on unseen content
Related research and updatesSynopsis
The authors introduce LUMOS, a diagnostic framework that uses OLMo 2 with its fully transparent training corpus to move LLM parametric-knowledge analysis from output-only inference to tracing the causal chain from training-data exposure to behavioral output, finding that models internally encode rare facts with high separability (84%) yet express them behaviorally only 54% of the time (a retrieval gap that narrows with scale), while self-reflection is reliable on trained content (83%) but drops to random-baseline levels (49%) on unseen content, a collapse that persists under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.
Figure 1 : A 2×2 taxonomy of model knowledge states defined by data exposure ( 𝒟 \mathcal{D} ) and behavioral accuracy ( A A ), enriched with parametric and decoding depth (Probabilistic Confidence ( P C PC ), Decisional Consistency ( D C DC )).
arXivInterpretation
It proposes LUMOS, a diagnostic framework that anchors parametric-knowledge assessment in verified training-data exposure rather than inferring what a model knows from its outputs alone. Prior analyses were largely output-centric and could not distinguish genuine generalization from rote memorization; LUMOS leverages OLMo 2's fully transparent training corpus to trace the causal chain from training-data exposure to behavioral output. The framework is built on OLMo 2 and its fully transparent training corpus, and the abstract reports quantitative results for internal separability, behavioral expression, and self-reflection accuracy.
Models internally encode rare facts with high separability (84%) but express them behaviorally only 54%, revealing a retrieval gap that narrows with scale. By measuring internal encoding and behavioral expression separately, the work turns the gap between knowledge presence and knowledge availability into a verifiable quantity rather than speculation. The abstract reports the specific values 84% and 54% and states that the retrieval gap narrows with scale.
When asked to self-reflect on their own answers, models perform reliably on trained content (83%) but fall to random-baseline levels (49%) on unseen content. It ties self-reflection reliability directly to training-data exposure, showing that reliability depends on whether the content was trained on. The abstract reports the accuracy values 83% and 49% and explicitly describes 49% as random-baseline levels.
Chain-of-thought prompting does not repair this collapse; it inflates confidence signals rather than improving calibration. This indicates the failure mode is not simply bypassed by prompting strategy, since prompting mainly adds confidence. The abstract states that the collapse persists even under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.
Perspective
The work targets LLM parametric-knowledge evaluation and applies to models whose training corpora are fully traceable, such as OLMo 2, letting researchers compare internal encoding, behavioral expression, and self-reflection reliability along a single causal chain; its conclusions are meant to show that incorporating the training-data axis turns judgments about what a model knows from speculation into verifiable claims, and it offers a framework for subsequent stratified analyses by scale and by whether content was trained on.
The currently visible text is only the abstract, lacking experimental detail, sample composition, and statistics, so the specific tasks, data splits, and evaluation protocols behind the 84%, 54%, 83%, and 49% values remain open questions; the exact trend by which the retrieval gap narrows with scale, how confidence inflation under chain-of-thought prompting is measured, and how the framework would be applied to models with opaque training corpora all need confirmation in the full text.
