Nature news feature headline: The Einstein test: what happens when AI tries to rediscover relativity?
The Nature feature. Illustration: Acapulco Studio.

On 9 September 2026, Nature published a news feature by Philip Ball. It follows one question: could a language model, given only what was known before a historical cut-off, arrive at an insight on the order of general relativity? Demis Hassabis proposed the test at the India AI Summit in New Delhi in February, suggesting a 1911 cut-off and calling it a good test for AGI. Owain Evans had asked something similar in a December 2024 talk about vintage models, and Ido Kaminer at Technion, who co-authored a July preprint asking whether AI can follow in Einstein's footsteps, argues the jump is reachable but not on the principles today's models are built from.

The feature's most useful move is naming the kind of reasoning involved. Tom Zahavy, a DeepMind researcher, argued in a January position paper titled LLMs can't jump that Einstein's advance was abductive: it invented a cause for a singular phenomenon. Jacob Andreas at MIT makes the contrast concrete, saying of Schrödinger writing down his wave equation that his mental process almost certainly did not look like generating theory-like statements at random and checking which matched observation. Thomas Kuhn called the result a paradigm shift, and shifts arrive where data are thin.

In an ICML paper, Vafa, Chang, Rambachan, and Mullainathan gave an orbital mechanics foundation model synthetic data from planetary systems that obey Newtonian mechanics, and it never recovered the law of gravitation. It produced a different law for each system, each wrong in its own way. Kepler's few observations gave Newton a gravitational law that carried over to every other pair of masses. Andreas adds a second gap: the models have no good notion of what an interesting mathematical theorem is, and their sense of what matters is where they are furthest behind human experts.

In May, an OpenAI chatbot disproved an 80-year-old conjecture of Paul Erdős, which Mullainathan called a genuine conceptual advance, though the disproof assembled ideas already in the literature. Michael Hla tried the historical version in March, training a model he called Machina Mirabilis on pre-1900 text. Shown the photoelectric effect, it said light "breaks up into a multitude of distinct impulses," which Hla read as a hint of quanta, but it failed in most cases and at points was parroting words that seem plausible.

The historical constraint turned out to be leaky. Nick Levine and colleagues trained a model on material through 1930, a date chosen because works published that year entered the US public domain in 2026, and it answered questions about Franklin D. Roosevelt's administration. At the University of Zurich, the Ranke-4B family uses cut-offs at 1913, 1929, 1933, 1939, and 1946, where team member Daniel Göttlich, now at ETH Zurich, said the target is not genius so much as sparks of genius.

Andreas puts the obstacle most plainly. Nothing stops a model from producing general relativity as one output among many bogus theories about the same data, and the difficulty is telling them apart. In mathematics each step is checkable. In physics it is not. Kaminer's group frames the same wall from the other side: major discoveries have often depended on persistent obsession, holding on to an outlier idea for years, and outliers are what these models average away. Their example is Planck, who in 1900 would not accept that his formula for black-body radiation gave results close enough to the measured spectra, and postulated quantized energy. Special relativity, on their account, had even less empirical pressure behind it, contrary to the common view about failed ether experiments.

Mullainathan says LLMs are not the right object by themselves, and that we need models trained on the physical world and on manipulating mathematical objects, neither of which is keeping pace. Kaminer wants these systems optimized for more than predictive accuracy, rewarding simplicity and explanatory reach. Against both sits Rich Sutton's bitter lesson: building in how we think we think, he argues, does not work in the long run, and only more compute and more training data move the number. AlphaFold sits awkwardly in the middle, predicting folded structures without solving how proteins fold.

Andreas expects the first results to be modest, in domains where we already know what the important problem is and where the answer can be borrowed from another field. He imagines a researcher thinking: damn, I wish I had done a second PhD so I could have seen that. Then he contrasts it with general relativity, where as he puts it the reaction was simply, where did that come from.

Editorial

The Einstein test is usually described as a test of creativity. What this reporting actually shows is two missing graders, and they stack.

The first is which answer to believe. A model that emits general relativity as one plausible output among many bogus theories about the same data has demonstrated almost nothing, because picking out the right one requires already knowing which is right. Physics has no cheap checker: to test a candidate law you run it against the solar system or build the apparatus. Mathematics does, since a proof step can be verified in seconds by a machine. That asymmetry, more than any difference in depth of thought, is the most economical explanation for why these systems look so much more impressive in one field than the other.

The second is which problem to work on. Andreas's remark about interesting theorems belongs next to the orbital result. A model that fits a different wrong law to each planetary system is not merely failing at physics. It has no way to register that the question worth asking is what all the systems share, which is precisely the sense human researchers carry.

Kaminer's obsession point follows from the same mechanism. Averaging is what a sequence model does. The failure mode is structural rather than incidental: the anomaly that matters is the one that gets smoothed away, and scaling does not change what the objective rewards. Planck, on their reading, is the type of case, where a refusal to settle produced quantized energy.

This is where the live disagreement sits, narrower than the framing suggests. Kaminer wants objectives that pay for simplicity and explanatory reach. Sutton's bitter lesson says built-in intuitions lose to compute and data in the end. AlphaFold is evidence for both readings at once, having beaten the folding problem by refusing to model the folding, which is either the sidestep Sutton predicts or the explanatory objective Kaminer is asking for.

The experiment Levine described is the one worth running, precisely because it does not require the obsession question to be settled first. Train a model to January, ask what happens in June, and score the answer against what happened. That turns a question about genius into a forecast with a settlement date. Until something of that shape runs, a positive result from a vintage model stays contestable on contamination alone, and a negative result describes a small model starved of text. The test is well posed. It has no scorer yet.