Public articles linked to the same research event.
arXiv The authors introduce LUMOS, a diagnostic framework that uses OLMo 2 with its fully transparent training corpus to move LLM parametric-knowledge analysis from output-only inference to tracing the causal chain from training-data exposure to behavioral output, finding that models internally encode rare facts with high separability (84%) yet express them behaviorally only 54% of the time (a retrieval gap that narrows with scale), while self-reflection is reliable on trained content (83%) but drops to random-baseline levels (49%) on unseen content, a collapse that persists under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.
The authors introduce LUMOS, a diagnostic framework that uses OLMo 2 with its fully transparent training corpus to move LLM parametric-knowledge analysis from output-only inference to tracing the causal chain from training-data exposure to behavioral output, finding that models internally encode rare facts with high separability (84%) yet express them behaviorally only 54% of the time (a retrieval gap that narrows with scale), while self-reflection is reliable on trained content (83%) but drops to random-baseline levels (49%) on unseen content, a collapse that persists under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.
The authors introduce LUMOS, a diagnostic framework that uses OLMo 2 with its fully transparent training corpus to move LLM parametric-knowledge analysis from output-only inference to tracing the causal chain from training-data exposure to behavioral output, finding that models internally encode rare facts with high separability (84%) yet express them behaviorally only 54% of the time (a retrieval gap that narrows with scale), while self-reflection is reliable on trained content (83%) but drops to random-baseline levels (49%) on unseen content, a collapse that persists under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.
The authors introduce LUMOS, a diagnostic framework that uses OLMo 2 with its fully transparent training corpus to move LLM parametric-knowledge analysis from output-only inference to tracing the causal chain from training-data exposure to behavioral output, finding that models internally encode rare facts with high separability (84%) yet express them behaviorally only 54% of the time (a retrieval gap that narrows with scale), while self-reflection is reliable on trained content (83%) but drops to random-baseline levels (49%) on unseen content, a collapse that persists under chain-of-thought prompting, which inflates confidence signals rather than improving calibration.