Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

LUMOS traces parametric knowledge with OLMo 2's transparent corpus: rare facts are 84% separable internally but only 54% expressed behaviorally, and self-reflection falls to 49% on unseen content

The authors introduce LUMOS, a diagnostic framework that uses OLMo 2 with its fully transparent training corpus to move LLM parametric-knowledge analysis from output-only inference to tracing the causal chain from training-data exposure to behavioral output, finding that models internally encode rare facts with high separability (84%) yet express them behaviorally only 54% of the time (a retrieval gap that narrows with scale), while self-reflection is reliable on trained content (83%) but drops to random-baseline levels (49%) on unseen content, a collapse that persists under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.