BootLoops 1.0 is an open-source toolkit that helps language models carry out precise scientific calculations with checks built into the workflow. Developed by Harvard physicist Matthew Schwartz with Claude, it is designed to work with different models. Its aim is not to make every AI answer trustworthy by default, but to make specific results easier to test.
11
2
How BootLoops checks calculations
The toolkit brings scientific software together with protocols that define what counts as a checked result. These include comparing an answer against an independent computational route at points not used to fit it, and using a positive control to demonstrate that a check can detect a bad result. Its guidance also addresses provenance, so a tool that contributed to fitting an answer does not simply certify that same answer.
2
11
That approach suits tasks with clearly defined outputs, such as mathematical-physics calculations. It can provide evidence that a result passed specified tests; it does not establish that every model-generated claim is correct or that the research question itself is worthwhile.
2
5
Results reported with Claude
Schwartz’s team reported calculating 30 mathematical-physics integrals, including 15 results described as new, among them elliptic Feynman integrals. The work also included a solution to what Schwartz calls Watson’s “final problem,” a calculation associated with mathematician George Watson’s earlier work.
14
16
4
The broader effort was reported to have produced 36 manuscripts with 19 coauthors across 18 fields, including ecology and population genetics. These are reported research outputs, not evidence that every result has completed independent verification: Schwartz said several results were still being checked.
3
15
4
Why human expertise still matters
Schwartz uses “impedance mismatch” to describe the poor fit between what scientists need and what current language models do well. Models can be useful for bounded, computational tasks where the problem and checks are explicit. But scientific work also requires deciding which questions matter and interpreting whether an answer has real significance.
1
5
In Schwartz’s account, Claude made connections across disciplines that could be technically correct but scientifically unremarkable. Domain experts helped steer the work toward questions that mattered in those fields. The distinction is important: computational verification can help establish whether a calculation passes its checks, while scientific judgment helps determine what to calculate and what the result means.
5
What BootLoops does—and does not—establish
BootLoops offers a way to structure AI-assisted work around testable tasks, independent checks, and human expertise. Its reported results show the potential of that approach, but they should not be read as a blanket guarantee of correctness or as a substitute for further verification. A checked calculation alone is not enough to justify a high-stakes decision; the question, assumptions, and consequences still need human scrutiny.
2
4
5