Skip to main content
Back to timeline
AnthropicSource publication:

Von Hippel's nine-loop bounty: Anthropic's Claude computed the six-particle nine-loop amplitude for about $100 of compute, and Dixon independently validated it

Synopsis

Physicist and science writer Matt von Hippel publicly challenged AI companies to solve a hard scattering-amplitude problem on an academic budget; Anthropic's Liam Fitzpatrick and Siddharth Mishra-Sharma used Fable 5.1 inside the Claude Science platform, gave Claude a single problem statement and then mostly just told it to keep going, and Claude computed the six-particle (hexagon) nine-loop amplitude in planar N=4 super Yang-Mills two different ways, via the original bootstrap and via the indirect form-factor route, with the bootstrap portion costing about $100, equivalent to running 96 CPUs for a week; Lance Dixon of SLAC then independently validated the result, and Song He's group at the Chinese Academy of Sciences had also computed the symbol piece of the nine-loop amplitude with GPT-6

AI-generated editorial illustration: Yes, Claude can do Nine Loops

Interpretation

Claude directly computed the six-particle nine-loop MHV amplitude in planar N=4 super Yang-Mills, a frontier calculation no one in the subfield had completed. The prior record was the eight-loop amplitude obtained indirectly by Lance Dixon and collaborators in 2023 via the form factor and a symmetry called antipodal duality; Dixon himself thought doing the amplitude directly at nine loops would be too hard. The result was independently validated by Dixon, largely by working back from the nine-loop form factor, a validation effort he says took about two weeks; the text reports no error bars or statistical measures, so validation is expert human checking rather than an automated benchmark.

The computation fit inside an academic budget: the bootstrap route cost roughly $100, corresponding to 96 CPUs for a week, and either approach would have cost an end user around one or two thousand dollars. Von Hippel's challenge was explicitly framed around resources an academic has access to rather than millions of dollars of compute, and the outcome suggests the expected computational barrier was not the real obstacle. The cost figures come from the two Anthropic physicists describing their own run to von Hippel; there is no third-party billing record or independent audit in the text.

The run was nearly unsupervised: researchers gave one problem statement and then mostly just kept it going with messages like keeping at it and giving updates every four to six hours, and Claude wrote all the code from scratch and completed both routes. Von Hippel contrasts this with March, when AI did physics projects like a student, with smaller tasks, heavy hand-holding and mistakes; this was a frontier calculation normally tackled by top experts. Based on the prompt text quoted in the post and the Anthropic researchers' account; the text states the author does not know how many internal mistakes Claude made, only that the harness reached the end without an outside collaborator's input.

Dixon notes that Claude used the methods his collaborators developed over the years and presented the solution in the format they had already set up, so validating Claude's result also validates their earlier work. This frames the AI contribution here as execution and engineering organization rather than new physical principles; Dixon says the more soul-searching moment will come when models produce new physical insights before humans do. This is the first-person judgment of the validator; the text offers no independent controlled comparison of methods.

Perspective

The work targets the toy model N=4 super Yang-Mills, used in amplitudeology to hone techniques rather than to describe the real world; the author states plainly that the theory is not used to explain dark matter or anything real. Its most direct audience is amplitudeologists working with bootstrap and form-factor methods, and teams wanting to gauge how AI science platforms perform on long, fragile computational pipelines. The author suggests people in real-world amplitude calculations should also check whether AI science harnesses can one-shot frontier calculations there, and should have a plan for checking results. For a broader reader, the piece offers a concrete case: on a well-defined problem with known methods and a compute bottleneck, an AI can run the whole job autonomously within an affordable budget.

The author admits he did not get the answer he wanted: he hoped to see AI break a computational limit with an unexpected new method, but Claude used known methods with more compute and better software engineering, leading him to conclude he was simply naive about where the limit was. He is also unsure how far this generalizes, since toy models are studied by small sub-communities while real-world amplitude calculations are a wider, more competitive field that may have less low-hanging fruit. Dixon stresses that the computational recipe is extremely fragile, crashing like a failed souffle if anything goes wrong, and the text does not say how many internal failures and debugging cycles Claude went through. Cost figures, prompts and run duration all come from the parties involved, and Song He's group is described only as having computed the symbol piece, with its completeness and publication status not stated in the text.

Sources