AlphaGo team member Thore Graepel leaves DeepMind, arguing AlphaGo-style search—not LLM chain of thought—is the path to genuine machine reasoning
Synopsis
Using move 37 from the 2016 Seoul match, AlphaGo core member Thore Graepel argues the move came from AlphaGo's search machinery rather than intuition, contends that today's large language models—relying on next-token prediction and chain of thought—do not reason in a scientist's sense, and proposes borrowing AlphaGo's game-tree structure so that general reasoning systems maintain an explicit, inspectable epistemic state updated by evidence, with an independent component judging each step by how much it actually resolves uncertainty.
Interpretation
The article reinterprets AlphaGo's move 37: the policy network saw nothing special in it, rating it at roughly a one in 10,000 chance of being played by an expert human, and what actually selected it was the search machinery, which explicitly constructed and searched a game tree with thousands of branches weighing the future consequences of candidate moves. The popular narrative treats move 37 as a flash of pure machine intuition; the author, a member of the AlphaGo team, says this is a misunderstanding and that the creativity came from reasoning (search), not intuition. A first-hand participant account and mechanism description, giving the order of magnitude of the policy network's probability and the scale of the game tree, but no new experimental data or controls.
The article maps AlphaGo onto Kahneman's dual-system theory: the networks supply hunches (this move looks promising, this position looks won) and the search supplies deliberation, testing those hunches against the moves and countermoves that would follow, with neither half working alone. It aligns a machine architecture with the behavioral-science two-system framework, offering a concrete machine instance of 'intuition plus deliberation'. A conceptual analogy and mechanism description drawing on an existing theoretical framework, not a new empirical result.
The article argues today's large language models are essentially next-token prediction, amounting to System 1 in action; chain of thought has produced real gains, above all in mathematics and coding, but its intermediate reasoning is still produced by the same next-token prediction process iterated for longer, introducing no genuinely separate reasoning mechanism. It separates 'chain of thought equals reasoning' into 'a longer run of the same process' versus 'an independent reasoning mechanism', and names three specific gaps: no explicit, persistent, inspectable epistemic state; knowledge and the manipulation of knowledge interwoven in the network's weights; and chains of thought often concocted after the fact. Primarily an architectural argument; it cites that 'research has demonstrated' bots often concoct chains of thought after the fact, but the text gives no specific study, sample, or data.
The article proposes an alternative: just as AlphaGo maintains a game tree, a general reasoning system should maintain an epistemic state representing what it holds as settled, what it doubts, what it has ruled out, and which questions stay open, with reasoning understood as a sequence of moves that change that state (deducing consequences, breaking problems into parts, deciding what question to ask, calculation to perform, or experiment to run next), and with an independent part of the system evaluating each move by how much it actually resolves uncertainty, updating beliefs only when backed by evidence. It generalizes AlphaGo's game-tree data structure into an epistemic state and belief-revision mechanism for the open world, emphasizing auditability—what the author calls 'the scientific method on steroids'. A position piece and set of design principles; the author acknowledges open-world reasoning is harder than board games—the state of affairs is only partially known, the action set is large and variable, and consequences are stochastic or unknown—and no prototype system or evaluation results are presented.
Perspective
The article addresses readers concerned with AI trustworthiness and scientific discovery, especially practitioners in high-stakes settings such as medicine, engineering, and research. It offers a design direction rather than a finished system: generalizing AlphaGo's game-tree idea into an epistemic state for general reasoning, making reasoning an auditable sequence of evidence, inference, and belief revision, and using large language models to suggest ways of tackling a problem, interact with tools via APIs or code, and help assess whether a claim is supported by available evidence. The author explicitly frames this for open-world problems while acknowledging that the open world is harder than board games, so the approach is best treated as a starting point for research and system design rather than a deployable solution.
Readers should note that the claim about chatbots often concocting chains of thought after the fact is supported only by 'research has demonstrated', with no specific study, sample, or data given, so its scope requires tracing the original literature. The proposed epistemic state and independent evaluation component come with no prototype, metrics, or comparison to existing methods, leaving feasibility, cost, and behavior under partial observability and stochastic consequences as open questions. In addition, the article defines 'reasoning' by a scientist's standard, which differs from looser engineering usage, so readers should be careful when citing it across contexts.
