If there’s one thing I’ve learned in the past week, it’s that people love quote round-ups. If there’s another thing I’ve learned, it’s that scientists have a lot of feelings about AI in science. With their powers combined…
AI is doing math
In the past few weeks, OpenAI announced 10 new (apparently impressive, according to mathematicians) results achieved by an internal model; Anthropic followed with a Claude-discovered Riemann-related result (not a proof, but a substantial improvement on the lower bound).
In light of these results, Fernando Borretti refutes some common arguments for why AI might not transform mathematics.
“Computers are superhuman chess players, yet we don’t care, and continue playing as normal. Why should mathematics be different? The main reason, I think, is that chess is self-contained: results from chess don’t help us understand the orbits of the planets or the binding of drugs to protein surfaces. But mathematics, famously, is the great dynamo of science, the best language and method for understanding the world. A machine that can replace a human mathematician, but better and faster and cheaper, is materially useful; a better chess engine is not.”
Noah Smith, while maintaining that the role of the mathematician will continue, argues that a specific form of heroic individual contribution is likely on its way out.
“We normally think of AI as something that replaces a single human worker, but I think that’s the wrong way to think of it. AI is really a new world-mind — a technology that takes the accumulated findings of individual humans (or robots, or other sensors) and integrates them into a general picture of the world… [E]ach time one of us reports a new scientific finding or expresses a new perspective, that adds to AI’s knowledge and understanding of the world… But unlike human society, AI doesn’t need heroic human geniuses. Its ability to understand complex ideas is not limited by the ability of a single human brain to apprehend, intuit, or communicate those ideas.”
Paata Ivanisvili recounts the pace of mathematical discovery from the front lines.
A small episode from the AI–math battlefield:
Aug 4: a paper appears on arXiv https://t.co/eJyFl3cfAG claiming partial progress on the Bourgain–Brezis Sobolev conjecture.
Same day: I ask AI, “can you try solving the full conjecture?”
Aug 6: AI says, “done.” (The proof attempt and all iterations are in my Overleaf history.) I start asking experts if they would like to check the argument.
Aug 8: another paper appears on arXiv https://t.co/stCwYd7ik7 claiming a full resolution of the conjecture (also with help from AI).
The pace is getting wild.
Tim Gowers explores the question of what aspects of mathematics, specifically, LLMs are good at, and the implications of potential answers.
“If it is true that current models are particularly good at finding examples, that is probably not because they have a particular affinity for existential statements, but more because the proof-discovery methods that are appropriate for finding certain kinds of examples play to the obvious strengths of LLMs: wide knowledge and the ability to explore many paths of the search tree that humans would judge to have a low probability of success.”
Terry Tao gives a lecture on mathematics in the age of AI, and the need for the field to take a more active role in examining its own role.
“What are the precise goals, objectives, and values of our mathematical community, and the enterprise of mathematical research? Not just the explicit goals that we communicate to the public (or to funding agencies), but also the implicit goals that we actually seek in practice?”
And Iris Shi writes a beautiful, difficult-to-summarize-but-very-worth reading meditation on the present and future of math as her PhD comes to an end.
“My physicist friend once asked me what the point of doing research was if someone like Terence Tao could have figured out everything in my dissertation in a tenth of the time. I answered by pointing out that Terence Tao didn’t. Terence Tao did not find a small open problem posited by my advisor and publish a bite sized result making incremental progress. He has only so much time and so many other fish to fry. And in his absence, I was given the chance to touch the edge of knowing and experience the unmatched exhilaration of discovering structure in the labyrinth of everything.”
Is AI doing other science?
There’s also been a flurry of activity in AI for non-math-science, with the first AI-designed virus, breakthroughs in cyclone forecasting,1 a shake-up at DeepMind, and a new Jeff Dean-led AI for science company.
To figure out what we can actually say about AI’s capabilities in drug discovery, an all-star team takes a careful look at the evidence and finds it wanting (if promising).
“The focus of AI in drug discovery must shift from doing what can be done - such as modelling data that is readily available, but that is unlikely to move the needle - to doing what should be done, even if this requires, for example, substantial data generation.”
“‘AI in drug discovery’ can mean two very different things: advanced, rapidly developed methods that lack thorough validation, as well as mature, established methods that have only just become practical to use at scale in production environments. What is common to both situations though is that the translation from ‘AI’ to ‘drug discovery’ has not yet been observed in an efficient and sufficiently durable manner to lead to clinical impact.”
Noah Olsman posits that many AI-for-bio efforts implicitly view biology as chess, when it is actually more like Kriegspiel.
“There is a variant of chess, called Kriegspiel, where each player can see only their own pieces. The game is played with three boards: two players sit back-to-back and move their own pieces while a referee keeps track of both players’ movements and mirrors them on a third board. The players go back and forth taking turns and, each time, the referee tells them whether that move is legal, and if they took a piece. The identity of the taken piece is only revealed if it’s a pawn.”
“Recently I have come to think that the differences between chess and Kriegspiel serve as a good metaphor for the gap between theory and practice in the life sciences. I entered biological research as a theorist who believed that enough mathematics could turn study of life into a game of chess: difficult, but ultimately legible. After spending years as an experimentalist, though, I have come to think that biology is much closer to Kriegspiel. We make moves (experiments) with only partial information, infer the hidden state of the board from sparse feedback (data), and often learn what mattered only after the consequences have unfolded, sometimes weeks or months later.”
Meanwhile Carlos Outeiral lays out the case that the gap in performance between AlphaFold 2 and AlphaFold 3 is because both rest on modeling evolution rather than physics, which works better for shape (AF2) than binding (AF3).
“The thesis of this essay is that the root cause of the slowdown is reliance on coevolution. The fundamental breakthrough in AlphaFold 2 was its ability to mine the evolutionary history of a protein to rapidly explore conformational space. Within a single protein, residues that touch in the folded structure (“contacts”) are constrained to evolve in concert… So if you align a protein’s sequence against millions of its relatives in genomic databases, these correlated changes capture a faint statistical shadow, left by evolution, of which residues sit close in space. … Put another way: the model appears to have learned the physics of biomolecular interactions, when it has largely learned protein genealogy.”
A team at DeepMind discusses both the promise of AI agents for speeding up scientific ideation, and the yet-unsolved gap in speeding up the validation of promising ideas.
“Most scientists have no shortage of ideas. The challenge is knowing which ones to pursue. This is where agents are starting to help. An agent can digest a field’s accessible literature, making connections across disciplines that no single researcher would have time to trace.”
“At the moment, however, the validation gap in most disciplines is widening, not closing. An agent can propose a novel genetic lead to reverse cellular ageing, but cannot say definitively whether it actually works. This explains why companies like Google DeepMind, Ginkgo Bioworks and Lila Sciences are investing in automated labs. But they only suit some fields, are expensive to build and are still early in development. And even automation cannot rush nature’s clock.”
Daphne Koller breaks down the misalignment between where AI-for-drug-discovery efforts focus and where the actual bottlenecks in drug discovery are:
“AI will certainly generate new and better molecules at an unprecedented rate, but will it generate drugs that unlock diseases for which there is currently no meaningful treatment? … [T]he biggest step functions in our ability to drug the undruggable have historically come not from better molecular design tools, but from expanding our repertoire of therapeutic modalities: first biologics, then siRNA and antisense oligonucleotides, then gene editing. Each new modality opened a class of targets that was simply inaccessible before.”
And the CRUX team, in a study led by Peter Kirgis, Sayash Kapoor, and Arvind Narayanan, look beyond AI agents’ ability to make progress on clean, verifiable tasks, and test whether they could conduct open-ended, independent research. For now, no.2
“Our key finding was that while agents could solve the engineering problems necessary to do the research, they failed to produce original research at the caliber of a top ML conference. Our main takeaways:
The agents lacked the judgment to identify when a problem was adequately solved….
The agents lacked awareness about the resources available to them and the timeline for the project…
The agents did not creatively respond to feedback about poor research design…
The agents did not effectively backtrack from unpromising approaches….
The agents did not follow concrete instructions....”
AI is reviewing (and submitting) science
Some aspects of the validation gap (at least in computational or theoretical fields) may be taken on by AI review. This possibility has caused two stirs on twitter in the past week or so, one when Refine announced its partnership with AEA and Econometric Society (discussed here), and another today based on this post from Itai Yanai.
Efforts to evaluate AI review tools often face the issue of mediocre ground truth (i.e. did a human or an LLM think the review was good overall? Did they think the individual comments were substantive?). Paul Litvak instead inserts errors and checks if tools can spot them (spoiler: they seem to do better than humans).
“Most systems work by ensembling, having the LLMs judge each other’s output. That, as we’ll see, is a good strategy, since different LLMs are surprisingly uncorrelated as far as the errors they can catch. Nonetheless, this isn’t satisfying since we don’t know whether even pooled models miss some categories of errors. Using human peer review as the gold standard, as any professor will tell you, is also unsatisfying, since human peer review varies widely in quality. Of Claude’s suggestions, one stood out — inserting errors into papers and seeing whether a given AI system can spot them.”
Beyond reviewing papers, advances in AI-based reproduction and replication of data+code continue apace, this time with a discipline-agnostic tool from Liu, Tjiaranata, and Tan.
“We present VERITAS, the first end-to-end domain-agnostic replication framework built around CLI coding agents. VERITAS supports flexible input modes, extracts the paper’s claims independently of the original authors, and withholds those claims from the replicating agent to mitigate leakage. The pipeline spans claim extraction, replication, and verification, and aggregates the per-claim verdicts into an importance-weighted Replication Score for the paper.”
But what about replication of findings in new data? From an ambitious new benchmarking effort:
“Several benchmarks have been proposed to develop and evaluate LLM agents on research reproduction tasks, including CORE-Bench, PaperBench, and REPRO-Bench… Existing benchmarks operate under the assumption that the new data sample is readily available to the agent. There are no benchmarks designed to assess agents’ capability to replicate research claims in a setting where a new data sample must be retrieved in advance... [W]e introduce ReplicatorBench, a benchmark for replicating published research claims in social and behavioral sciences.”
Let’s close on a would you rather: (a) have an LLM review your conference submissions, or (b) review dozens of LLM-authored conference submissions? I don’t know for sure which they’d pick, but I can say that Caleb Robinson and Isaac Corley did not enjoy the conference review circuit this summer.
“Between the two of us, we reviewed 22 paper submissions this summer, spread across NeurIPS, WACV, and TerraBytes (a geospatial workshop at ECCV). Fifteen of the 22 (68%) contained entirely fabricated citations, fabricated author lists for existing papers, and/or were clearly LLM-generated (e.g. hallucinated technical jargon, nonsensical writing, irrelevant citations). This is called being in the ‘slop trenches.’”
There was a funny cognitive offloading tweet about this one that I wanted to link but can’t find. If that was you, or if you know it, let me know!
To be fair to the agents, it doesn’t look like the team tried prompting “it’s enough of partial results. let’s finish with a complete breakthrough study.”




