LANTERN used Qwen3-32B hidden states to rank 50 million OEIS sequence pairs and returned 62 verified relations in under 8 hours, four of them absent from OEIS and a targeted literature search
Synopsis
LANTERN trains a linear classifier on Qwen3-32B activations to rank 50 million candidate relations among the 10,000 most-referenced OEIS sequences, then applies staged filtering, hypothesis generation, executable verification and analytical checking to produce 62 verified relations without an existing OEIS cross-reference; a content screen retained 13, nine of them informative or insightful, four not found in the OEIS or a targeted literature search, with the end-to-end process taking under 8 hours.
Interpretation
A pretrained language model already encodes previously undocumented relations between mathematical objects, and that knowledge can be used directly to rank candidate relations. Most prior AI discovery systems rely on human-provided objectives, statements already written in the literature, knowledge-graph topology, or the model's own generation; this work instead ranks every pair of objects in a knowledge base by what the model's hidden states say about them. A linear classifier was trained on OEIS substantive cross-references as labels and scored all 50 million pairs of the 10,000 most-referenced entries; in controls, a surface-text scorer was at chance (AUC around 0.5) on 793 held-out links, whereas activation-based classifiers reached AUC in roughly the 0.6 to 0.7 range across three model sizes.
LANTERN is a cost-effective, efficient discovery pipeline that chains ranking, filtering, hypothesis generation and verification end to end. The work provides a reusable four-stage process: classifier ranking, a cheap filter, a tool-equipped agent proposing hypotheses, and sandbox numerical verification plus analytical checking, with per-stage timings reported. The main configuration narrowed 50 million pairs to 500 ranked candidates, 118 passed the inexpensive filter, 44 received a hypothesis, and all 44 passed sandbox verification and analytical checking; classifier training and scoring of 50 million pairs took 8 minutes, and the whole process including labelling, embedding, filtering and the agent took under 8 hours.
On the OEIS the pipeline produced 62 technically verified relations; a content screen retained 13, of which nine are informative or insightful and four were not found in the OEIS or a targeted literature search. These are pairs the OEIS does not cross-reference; two cross-domain bridges (one- and two-dimensional cellular automata, and a continued fraction with 5-core partitions) are labelled N2/S2, while two N2/S1 relations supply computational routes. The 62 statements are the union of three configurations (44 from the main configuration, 3 and 15 from two variants); the content screen judged each item on exactness, specificity to the pair, and use of substantive structure from both objects, retaining 13; novelty was assessed against the OEIS and a targeted literature search, and the authors state that N2 does not establish priority.
Controls separately rule out surface text, page memorization and the training labels as alternative explanations. The work trains a text scorer on the same labels and negatives, runs post-cutoff tests on cross-references added after the training snapshot and on entries created in 2026, and adds a label-free direct question as a readout. On 2,379 entries created in 2026 and their 5,557 cross-references to older entries, a classifier trained only on cross-references among older entries reached AUC in roughly the 0.7 range, while the text scorer dropped from roughly 0.8 to roughly 0.6; the label-free question and the classifier achieved nearly identical median percentiles on 74 substantive future pairs (88.5 and 88.0).
Perspective
The result is aimed at settings with a large structured knowledge base whose objects are well defined, whose relations are documented, and whose hypotheses are relatively cheap to verify; the OEIS satisfies all three and serves as the test bed. The method itself can transfer to other knowledge bases with explicit links and executable verification, and the authors conclude that pretrained representations can serve as a practical guide for discovering unrecorded relations. What a reader can reuse directly is the four-stage process and the control design: rank candidates with hidden states, compress the queue with a cheap filter, generate executable hypotheses with a tool-equipped agent, then check numerically and analytically. The authors also note that the 62 statements are the union of three configurations rather than a single run, that 13 were retained by the content screen, that nine are S1/S2, and that four were not found in the OEIS or a targeted literature search.
The derivations were produced by a language model and have not been checked by a mathematician, so the mathematical status of the 62 relations still needs human review; N2 only means the connection was not found in the OEIS or a targeted literature search, and the authors state it does not establish priority and do not claim the underlying identities are new theorems. The OEIS is a highly structured test bed, the corpus is biased toward frequently mentioned sequences, and only one family of pretrained models was evaluated, so applicability to less formal objects or other knowledge bases remains open. Control conclusions also vary with the evaluation set: on links added after the training snapshot the scorers do not separate, whereas on the topic-matched test the activation classifiers outperform the text scorer; Appendix D further notes that the 74 future pairs correspond to about 45 distinct facts, more than half of them of two templates, which affects how the scorers' generalization should be read. In addition, for 20 statements in Appendix A the content screen read only the entries' formula and comment lines rather than the full entries, a scope limit that affects how strictly the screening should be judged.
