Skip to main content
Back to timeline
arXivSource publication:

SourceLearn turns repeated use of one authoritative source into cumulative source-specific competence, topping 13 of 15 settings

Related research and updates

Synopsis

The work formulates source learning, represents reusable understanding of a persistent authoritative source as a persistent revisable source model, and refines it through Self-Directed Source Learning (an Inspect-Study-Consolidate cycle with adaptive Deepen or Connect actions) and Task-Guided Source Learning (failure-guided local refinement plus cross-task representation-policy recalibration), with persistent updates always reconstructed from the authoritative source; across five benchmarks and three LLM backends, SourceLearn is best in 13 of 15 settings, improving over Hybrid RAG by 14.3, 4.9, and 13.4 points on average, with gains up to 22.6 points.

AI-generated editorial illustration: From Knowledge Access to Source Learning: Developing Source-Specific Competence

Interpretation

The paper formulates source learning as developing reusable source-specific competence over a persistent authoritative source, represented by an explicit, persistent, and revisable source model that complements direct source access while the source remains the factual authority. Prior work separately improved how source content is accessed and organized (retrieval, long-context modeling, structured source representations) or preserved knowledge from prior interactions (agent memory), treating repeated use of the same source as repeated access; here the object of learning is understanding of the source itself. Presented through the problem formulation (Eqs. 1-3) and the design of the source model, entity representations, and activation budget (Eqs. 6-7); this is a conceptual and representational contribution whose value is supported indirectly by the experiments.

SourceLearn builds and progressively refines the source model with two complementary mechanisms: Self-Directed Source Learning, an Inspect-Study-Consolidate cycle that uses the current model to guide renewed reading and an adaptive planner to select Deepen or Connect actions, and Task-Guided Source Learning, which uses downstream task experience to reveal local representational gaps and recurring representational needs. Both mechanisms share one grounded-reconstruction principle: learning signals determine what should be reconsidered, while persistent updates must be reconstructed from the authoritative source, and observations are never copied directly into persistent memory; task-guided learning further separates failure-guided local refinement from cross-task representation-policy recalibration. The method is specified in Eqs. 4-14 with prompt details in Appendix C; ablations show that removing either mechanism or either pathway lowers accuracy on all three QA benchmarks, with the initial source model lowest.

Across five benchmarks (MultiDoc2Dial, NarrativeQA, SWE-QA, APIBench, AppWorld) and three backends (GPT-5.6-Luna, gpt-oss-120b, DeepSeek-V4.1-Flash), SourceLearn is best in 13 of 15 settings and second in one more, improving over Hybrid RAG by 14.3, 4.9, and 13.4 points on average, with gains up to 22.6 points. Counterparts include task-local retrieval (Hybrid RAG), static source representations (RAPTOR, HippoRAG 2), and experience-based memory (AWM, which learns from the same guidance tasks); the paper reports that the latter three improve more unevenly across benchmarks. The main table spans document, code, and API sources; guidance and test tasks are disjoint, QA is judged by GPT-5.6-Luna, and APIBench and AppWorld use their official evaluators; under matched context budgets SourceLearn remains best on three of four sources.

Source learning goes beyond compression: from the initial model to the self-directed and then task-guided model, representation shifts from isolated look-up content toward structured rules and procedures, and units with explicit applicability conditions rise from roughly 30% to 65-82%; the share of source claims required by test questions that are represented rises from 23.2% to 40.0%, and coverage is associated with downstream success (67.5% accuracy with none represented versus 86.6% with all represented). This addresses the concern that the source model is merely a compressed summary, using representation composition, condition explicitness, and required-claim coverage analyses to show that learning changes what is represented. The analyses rest on category groupings of source-model units and claim-coverage statistics; the paper also reports a grounding audit with few direct contradictions and 93-100% of dropped claims preserved after rewriting, while some task-guided units on document and code sources are only partially supported or unsupported.

Perspective

The result targets persistent, authoritative, and relatively stable sources together with repeated same-source task sequences; it applies to LLM agent systems that must complete successive tasks over one document collection, code repository, or API documentation set. It enables later work to combine an activated source understanding with directly retrieved source evidence at task time, to retain usable source understanding when retrieval is incomplete, and to convert task experience into improvements of the source representation rather than retention of task answers. The paper reports that with no retrieved evidence the source model alone performs comparably to Hybrid RAG with all eight retrieved elements, and that removing the gold document lowers Hybrid RAG by about 25 points but SourceLearn by only about 8; under matched context budgets SourceLearn remains best on three of four sources.

The per-cell numbers of the main table are not shown in the supplied material, so per-benchmark and per-backend gaps can only be judged from the prose summary; the grounding audit reports partially supported or unsupported task-guided units on document and code sources, which is why the source model is positioned as accumulated understanding rather than a replacement for the source. Extending self-directed learning to four cycles produced no consistent gain beyond the first cycle, leaving the relation between optimal cycle count and source size or type open. The paper lists source learning over evolving, noisy, or conflicting sources as future work, so behavior in those settings remains to be examined.

Sources