Schwartz used BootLoops to have Claude compute 30 Feynman integrals end to end, 15 of them never computed before
Synopsis
In a guest post, theoretical physicist Matthew Schwartz describes building BootLoops, an open-source harness for exact calculations in quantitative science: rather than asking Claude to work like a human scientist, he looked for "Claude-shaped" problems, and Claude reproduced in about 20 minutes results whose code took him weeks to write, completed 30 Feynman integrals end to end (15 reproductions of known results by the new method and 15 never computed before), then carried the same class of integral methods into ecology and population genetics, advancing 36 manuscripts across 18 fields with 19 coauthors in three months.
Interpretation
The author proposes and practices a problem-selection principle he calls "Claude-shaped" problems: instead of treating current LLMs as human scientists, look for problems matched to their strengths, and encode that match in BootLoops, an open-source harness. Relative to earlier use of the model as a research assistant requiring sentence-by-sentence correction, he first identifies what the model does well (breadth of knowledge across domains, coding skill, leading-edge mathematics and statistics, parsing papers and data at machine speed) and selects problems accordingly. This is a first-person account based on the author's own experience, supported by the contrast between Claude Opus 4.5 as a research assistant whose every sentence had to be corrected and the later capability-matched approach; it is experiential rather than a controlled experiment.
On the semi-numerical S-matrix bootstrap, Claude ported methods scattered across Wolfram Language, C++, Python, Julia, and papers with no code into a common framework, generalized them to elliptic function classes, and completed 30 integrals end to end, of which 15 reproduce known results by the new method and 15 had never been computed before. The author states that only a handful of elliptic Feynman integrals had ever been computed and, to his or Claude's knowledge, none completely by the bootstrap; the obstacle was not that the methods would fail but that nobody had tried, because the needed expertise is distributed among many humans. The text offers a verifiability argument: the same numerics let anyone, expert or not, check the final answer against the integral to as many digits as they like by running two scripts; it also gives a time comparison, with Claude reproducing in about 20 minutes results whose code took the author weeks.
Carrying the same class of integral methods into other fields, the author and domain experts obtained concrete results: in ecology, solving Etienne's 2005 equation that nobody could solve at scale for 20 years, yielding a finding that on Barro Colorado Island in the Panama Canal the mix of tree species changes 4.5 times faster than neutral theory allows; in population genetics, analyzing 5.7 billion pairs of nearby mutations from the 1000 Genomes Project and finding evidence for gene conversion. The author stresses that these connections were at first technically correct but scientifically unremarkable, and became meaningful only after experts such as James O'Dwyer and Michael Desai redirected them toward questions their fields care about, for example subtracting the neutral prediction and studying the remainder, or turning to correlations between pairs of mutations on a single chromosome. The ecology result rests on computation applied to an existing equation and to data from a heavily studied forest; the population-genetics result rests on analysis of 5.7 billion pairs of nearby mutations; the author also notes these works are undergoing further exploration and verification.
The author describes an engineered way to run many projects at once: Claude Code sessions run in terminals on Google Cloud virtual machines linked to GitHub and Overleaf repos, with one session per project plus a master session that coordinates the others, allocates compute, and validates results, alongside separate sessions for writing, for creating repos and tool manuals, for writing and validating tool code, and for adversarial re-checking of results. Relative to single-conversation use, this layered session and background-subagent structure let him advance 36 manuscripts across 18 fields with 19 coauthors in three months, out of some 400 candidate problems, and let classifier blocks shut down a single agent rather than corrupt the whole session. This is the author's direct description of his own workflow, including concrete scale figures (36 manuscripts, 18 fields, 19 coauthors, three months, about 400 candidate problems), but no control condition or quantitative evaluation is provided.
Perspective
This work targets quantitative, checkable problems that can be decomposed into coding and algorithm development; the author states plainly that the current model cannot help him with deep conceptual questions, and that humans still own the conceptual and directional part. Settings where it applies include computations like the semi-numerical bootstrap that draw on mathematics, physics, and computer science no one person has mastered and whose results can be verified numerically; data-rich but theory-starved fields such as systems biology, which the author names; and domains with large public datasets where AI can surface hidden biology, since he notes thousands of genomic datasets sit in public archives. The author also positions BootLoops as an open-source harness meant to be used broadly in quantitative science, with whatever model the user prefers, without each user rebuilding everything from scratch.
The author repeatedly flags caution: outside his own field he found himself agreeing with Claude's judgments, so he brought in experts and found that in almost all cases Claude was technically correct but the result was not all that interesting until the expert helped steer. He states explicitly that some highlighted work is "undergoing further exploration and verification." In addition, several places in the text (for example "For example:", "Some additional highlights ... include:", and "Here are a few more failure modes I encountered, and tips for dealing with them:") are not followed by the promised content, so those specific cases and the failure-mode list cannot be learned from this article. He also notes that Claude has no sense of time, that its ETA estimates are always far too long or too short, that it defaults to grinding through long calculations rather than building a tool, that compaction causes it to lose context on long projects, and that the projects were compute- and token-intensive. Readers should still watch for independent replication and peer review of these results in their respective fields, and for how the credit-assignment and graduate-training questions he raises get resolved.
