50 Comments
russell
russell
27 days ago

So sorry, cleek.

GftNC
GftNC
27 days ago

Back to LLMs (all links removed). I thought this interesting:

Einstein, Churchill, and AI
What LLMs Can’t Do
Ian Leslie
Aug 15, 2026

I am not very fond of arguments over what AI can or cannot do, since they tend to devolve into arguments over definitions. For instance, it used to be said that AI can’t reason. Now that it’s solving maths problems which have stumped our brightest mathematicians for decades, the sceptics are saying ah, but that’s not what we meant by reasoning. On the other side, AI boosters address every capability gap with a “Not yet. The technology will obviously continue to improve until it solves whatever shortcoming you’re pointing to.” Well, maybe, but this is a frustratingly unfalsifiable claim which can be wheeled out to demolish any sceptical argument, a literal deus ex machina.

So I appreciated this new paper from a Google DeepMind researcher called Tom Zahavy, which is – unusually for a paper from an AI lab – all about what AI, at least in its current form, cannot do. Its title is “LLMs can’t jump” (Ron Shelton’s movie surely has one of the best and most generative titles of all time). In short, it’s about why LLMs are unlikely to make scientific breakthroughs on their own.

The paper argues that there is a fundamental limit to their capacity for invention. LLMs can ingest tons of data and identify the patterns and rules which govern a particular domain. They can also work their way through a series of logical steps to solve specific, highly complex problems. What they can’t do is make a creative “jump” – the sudden insight that enables a human scientist to go beyond the available data.

Zahavy gives what is possibly the most famous documented example of such a leap. In 1907 (the same year Pablo Picasso re-invented art), Albert Einstein was working at the patent office in Bern and also on a new theory of the universe which yoked together space and time. While pondering how “relativity” would work for someone speeding up or slowing down, he had what he later described as “the happiest thought of my life”.

Einstein imagined a man falling from the roof of a house, or standing in a free-falling elevator. What would that feel like? It would feel weightless. If he was holding an apple and let go of it, the apple would “float” alongside him. From the falling man’s point of view, then, it could be said there was no gravity in his pocket of space. That weightlessness, Einstein realised, would be exactly the same if the man were floating in deep space, far from any planet’s gravitational pull.

In a sense, then, falling, not standing, is the man’s natural state. When he stands, the earth is preventing him from falling. His weight is the earth pushing back. From there, Einstein eventually arrived at the idea that objects don’t fall because gravity is pulling them, as Newton believed, but because mass changes the shape of space and time, and objects move along the paths which that shape creates.

If you don’t fully understand this, that’s OK, neither do I. What matters, for our purposes, is that this epochal breakthrough was based on very little data. It’s not like early twentieth century astronomy was generating lots of conflicting information about gravity and acceleration and space. An 1907 LLM trained on Newtonian rules and fed data on planetary movements would have received barely a squeak of an error signal. There was a tiny unexplained wobble in Mercury’s orbit, which most scientists explained by reference to an unseen planet. Otherwise, Newton had it covered.

But Einstein was unsatisfied. Newtonian gravity and Maxwellian electromagnetism worked in completely different ways, and Einstein couldn’t accept that the universe ran on two separate rulebooks. The universe was one, he just knew it – and this led him to make his conceptual leap.

Back to LLMs. Zahavy notes that Einstein’s insight was rooted in physical experience – in knowing what it was like to be a body subject to gravity. It had to be felt because it could not yet be expressed in abstractions – in words or mathematical symbols. Newtonian science offered no logical explanation for his imaginary scenario. Einstein himself described his thought process like this: ”The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought.”

LLMs are built out of words and symbols. They don’t have physical experience of the world. They are cathedrals of pure abstraction. If Einstein had an LLM, it could have helped him work through the consequences of his discovery and to make predictions, which involved a huge amount of mathematical work. He would have appreciated that; he wasn’t a very enthusiastic mathematician. But the LLM couldn’t have made the foundational leap for him. No amount of pattern-matching and deduction would have got it there.

Zahavy points out that many scientific breakthroughs did not make sense in the light of available data. They had to be intuited or imagined into being. That in turn has depended, as unscientific as it sounds, on what the scientist wants to believe – on some kind of aesthetic or moral demand they are making of the world. Great scientists often have deep, supra-scientific convictions which operate like search heuristics, pushing them to look for explanations that can’t be spotted in the data.

He mentions Kepler’s neoplatonic belief in the centrality of the sun. We might add Faraday’s discovery of the electromagnetic field (then given mathematical form by Maxwell). That sprang, in part, from Faraday’s Christian belief that since nature was the orderly creation of God, apparently separate forces should exhibit underlying unity. Similarly, Einstein had a profound philosophical conviction that the universe is a unified whole. He wanted that to be true, which pushed him to discover that it is true.

To design an AI capable of invention, Zahavy says, technologists will need systems that “do not just simulate the world but hold strong beliefs or priors about how that world should be structured…”. Note that “should”. Can a machine have a passionate conviction or belief? Would that count as a form of “intelligence” or is that word inadequate to describe the thought process of an Einstein? Zahavy doesn’t explore those questions, but for now, this kind of cognition remains uniquely human.

C.P. Snow’s collection of biographical essays, Variety of Men, includes vivid portraits of Einstein and Churchill, both of whom he had met. He detected certain likenesses. Neither did very well at school, and both developed a deep and abiding conviction that they were different, with something unique to contribute to the world. Once settled on a belief or idea, both were “unbudgeable”; Einstein quietly, Churchill loudly.

In a brilliant passage on Churchill, Snow draws a sharp distinction between judgement and insight. In leadership terms, good judgement means being able to grasp the complexity of a situation and assess the trade-offs. Churchill, says Snow, was not good at this. In fact, most contemporaries agreed that his judgement was “seriously defective”. Here’s Snow:

Churchill had a very powerful mind, but a romantic and unquantitative one. If he thought about a course of action long enough, if he conceived it alone in his own inner consciousness and desired it passionately, he convinced himself that it must be possible. Then, with incomparable invention, eloquence, and high spirits, he set out to convince everyone else that it was not only possible but the only course of action open to man. Unfortunately, the brute facts of life were not always so malleable as his listeners.

Churchill’s obsessive pursuit of certain ideas led to blunders like Gallipoli and his vain attempts to prevent Indian independence. But it was this same obsessiveness which led to his greatest achievement. When all the smart people with good judgment in the British state were proposing that Hitler be accommodated, Churchill alone thought otherwise. Snow again:

Judgement is a fine thing, but it is not all that uncommon. Deep insight is much rarer. Churchill had flashes of that kind of insight, dug up from his own nature, independent of influences, owing nothing to anyone outside himself. Sometimes it was a better guide than judgement. In the ultimate crisis, when he came to power, there were times when judgment itself could, though it did not need to, become a source of weakness. When Hitler came to power, Churchill did not use judgment, but one of his deep insights: this was absolute danger. There was no easy way round.

You often hear it said that “judgement” is where humans will retain an edge over machines. But “insight” might be an even wider moat, and if the AI doomsters are right then it will become very necessary at some point.

Yann LeCun, former chief AI scientist at Meta, now an AI startup founder, was recently asked what he wanted his legacy to be. He said “Increasing the amount of intelligence in the world.” With more intelligence, he continued, there will be less human suffering, more rational decisions, more understanding of the world and the universe. But more intelligence, in the sense he means, wouldn’t have produced Einstein’s leap or Churchill’s insight.

Were Einstein and Churchill obviously more intelligent than their peers? No. Einstein was the greatest scientist of the twentieth century but not because he had a higher IQ than Planck or Bohr (he acknowledged that many of his peers were better at maths). Churchill’s visionary leadership wasn’t the product of rational analysis. Physics and politics are very different realms and Churchill and Einstein had very different kinds of mind, but they both had this ability to “dig up from their own nature” insights that nobody else had grasped, and which weren’t in the data.

nous
nous
27 days ago

Interesting piece from Leslie. Many “yes!” moments from me, but also a few “no, dammit!” moments where I thought he was missing something deeper, or anthropomorphizing in ways that misrepresented the actual situation:

Now that it’s solving maths problems which have stumped our brightest mathematicians for decades, the sceptics are saying ah, but that’s not what we meant by reasoning.

Well, it’s not. I’d argue that the LLM did not, in fact, solve any unsolved problem because it did not set out to solve the problem knowing that there was a gap in our collective knowledge, nor is is able to understand and decide that it had found a correct solution, and know how that solution could be applied. It has no world and no experience. The researchers set it to solving the problem and prompted the parameters. There was no moment of understanding or synthesis for the bot.

I’m reminded of the technicians that were attempting to find the cause of the anomalous signal in a radio telescope, who isolated every possible functional answer for how a telescope could generate this noise itself and ruled them all out, thus proving the presence of the background cosmic radiation predicted by the Big Bang theory.

That was not a discovery prompted by insight on their part. They had not set out to discover cosmic background radiation. They had tried, and failed, to eliminate a noise source that was interfering with the theoretically optimal operation of a radio telescope. The recognition and discovery happened when they submitted their results to scientists who could make the intuitive connection between their work and the theoretical existence of CBR. Without the scientists, they would have just had a radio telescope with an unknown noise problem.

Back to LLMs. Zahavy notes that Einstein’s insight was rooted in physical experience – in knowing what it was like to be a body subject to gravity. It had to be felt because it could not yet be expressed in abstractions – in words or mathematical symbols. Newtonian science offered no logical explanation for his imaginary scenario. Einstein himself described his thought process like this: ”The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought.”

LLMs are built out of words and symbols. They don’t have physical experience of the world.

Yes!

They are cathedrals of pure abstraction.

No! Or at least this is anthropomorphizing them and imagining an independence, and awareness, and an understanding that they do not have. They do not live in a world, and cannot therefore abstract anything from the world to intuit anything higher order.

Similarly, Einstein had a profound philosophical conviction that the universe is a unified whole. He wanted that to be true, which pushed him to discover that it is true.

No. Einstein discovered that his theory based on that conviction solved problems with the existing models, and gave us a means of generating new problems that we could use to test the validity of his conviction in other circumstances to see how well his convictions held. I think we have two different species of “truth” at work here.

To design an AI capable of invention, Zahavy says, technologists will need systems that “do not just simulate the world but hold strong beliefs or priors about how that world should be structured…”. Note that “should”. Can a machine have a passionate conviction or belief? Would that count as a form of “intelligence” or is that word inadequate to describe the thought process of an Einstein?

And just how, exactly, is the AI supposed to form these strong beliefs about a world that it does not live in in any meaningful way? What constitutes a “passion” in a binary assembly? Can a passion be prompted?

Judgement is a fine thing, but it is not all that uncommon. Deep insight is much rarer. Churchill had flashes of that kind of insight, dug up from his own nature,

Yes, in the sense that he is a self-directed being living in a world and forming judgments about that world, shaped by experience, that can be tested to shape his understanding of that world.

independent of influences, owing nothing to anyone outside himself.

No! Churchill, like the LLMs, is working from a dataset and is highly dependent on “influences,” in that his language and society are both acquired and maintained by collective consensus. He can indeed intuit conceptual jumps that did not exist before he makes them, but those insights are not independent of influences, they are just novel approaches that have the potential to reshape consensus. The insight is still predicated upon a world and a society that pre-exist, and co-exist, whose influences permeate his every thought.

wjca
wjca
27 days ago

LLM did not, in fact, solve any unsolved problem because it did not set out to solve the problem knowing that there was a gap in our collective knowledge, nor is is able to understand and decide that it had found a correct solution, and know how that solution could be applied. It has no world and no experience. The researchers set it to solving the problem and prompted the parameters. There was no moment of understanding or synthesis for the bot.

This is, I think, the core disagreement between the AI enthusiasts and those of us who are skeptics. The AI does interact with the natural world. All it can do, all it will ever be able to do without a complete redesign, is respond to specific queries from those who do.

Take an obvious (and frankly worrisome) hypothetical. Suppose an AI is fed a query: prove the Donald Trump is the greatest President ever. It will put together a line of plausible sounding bull to that end. Given a query substituting “worst” for “greatest”, it will produce equally plausible sounding bull.

Neither will be particularly constrained by the fact that whatever statistics it uses might be invented out of whole cloth. If it needs inflation to be low, it’s got plenty of sources (starting with Trump himself) who say that it is. But the AI has no reality check of shopping every week.

GftNC
GftNC
27 days ago

Sorry to hop about like this, but that’s the beauty of an open thread! I occasionally look at Doc Science’s bsky thread, and she links a Vogue article about the 5 books Emily Wilson says influenced her in her translation of the Odyssey. I was delighted to see that one of them was A Wizard of Earthsea by Ursula LeGuin, since I too am very keen on those books. This is what she says about it:

I’ve loved this book since I was a child. I’ve reread it many times. It is, like the Odyssey, a book about travel, identity, and a clever man’s journey away from one community and back to another. But primarily, I thought of LeGuin’s style, especially in that book, as one of many guiding lights as I considered and experimented with how to convey the vividness (enargeia in Greek) of Homeric poetry. Of course, LeGuin is writing prose, not metrical poetry, but there’s something both very beautiful, very plain/noble/rapid, and very Homeric about her style in this book. The sentences actually often have a beautiful rhythmicality to them: “This is a tale of the time before his fame,
before the songs were made.”

Pro Bono
Pro Bono
27 days ago

There seems to be some confusion, including in Zahavy’s paper, between LLMs and AI more generally. Zahavy himself must know what he’s talking about, but he obscures that when he extrapolates without explanation from the achievements of AlphaProof (not a LLM) to what LLMs ought to be able to do.

(I’m reasonably sure that the OpenAI model which disproved the Unit Distance Conjecture is not a LLM either, though they seem not to have said much about it.)

This distinction between LLMs and AI designed for a specific purpose is rather important to understanding what AI can do. For example, LLMs are very bad at chess, whereas AI (Stockfish is the leading program) is vastly better than even the strongest humans.

nous
nous
27 days ago

A bit more unpacking of why I bristle at Leslie saying that Einstein discovered that his conviction that the universe was a unified whole was true.

Here’s Sabine Hossenfelder explaining succinctly and elegantly how it is that we know that there are problems with Einstein’s theory:

https://youtu.be/Ov98y_DCvRY?si=gN6_Z72P8H7lrY7I

We have yet to get Relativity and Quantum Theory to play nice with each other. Both are very good at explaining how the universe acts on many levels, supported by our observations to our best currently capable degree of measurement. (Can we say “verified?” That seems just a bit too strong, where “supported” seems too weak.) One or both of them must be either incorrect, inconsistent, or unverifiable. We are currently unable to tell which of these is the valid conclusion.

I’m sure that Hossenfelder would say that Einstein’s theory is reliable within the known bounds, but we’ve also known since the 1930s that it’s not sufficient to explain all the things it attempts to make sense of.

wjca
wjca
26 days ago

One or both of them must be either incorrect, inconsistent, or unverifiable. 

I think the word you’re looking for is incomplete. Just as Newtonian mechanics is not incorrect, nor inconsistent, nor unverifiable. But it is incomplete.

For an enormous number of situations, Newtonian mechanics still works just fine; for some situations its incompleteness matters. Likewise, for a lot of purposes (e.g. chemistry), treating atoms as a collection of protons and neutrons surrounded by electrons is fine. You only care about the quarks that make up those one-time elementary particles in special situations (e.g. atomic physics).

That’s actually rather usual in science. We have an understanding. We find something which doesn’t fit. We develop a revised understanding. It’s an ongoing process of successive approximations. My personal suspicion is that we aren’t even close to an ultimate understanding.

nous
nous
26 days ago

wj – They are both incomplete for now, but if we could find a likely way forward, then either one or both of them would prove incorrect, one or both of them would prove applicable only in a more limited set of circumstances (inconsistent), or as seems the case with many of the proposed solutions for quantum gravity, the answer is beyond the scope of what we can observe and/or measure, and thus be unverifiable.

We can both be correct here.

wjca
wjca
26 days ago

Ah. I was taking unverifiable as “beyond the scope of what we will ever be able to observe and/or measure.” Whereas it sounds like you only meant “what we can currently observe and/or measure.”

nous
nous
26 days ago

In at least one of Hossenfelder’s videos she says that in order to measure fluctuations in Einstein’s space time continuum to test quantum gravity we’d have to be able to measure differences 10 to the 23rd power times smaller than an atom.

Godspeed my friend.

Sounds like “may never know” to me.

Hartmut
Hartmut
26 days ago

Hey, that’s still more than the Planck length (by an order or two of magnitude depending on how one defines ‘diameter of an atom’), so no need to poke black holes into the space time continuum.
If it was as -275°C, that would be another matter (and I don’t mean dark one despite it being difficult to shine a light on it).

Michael Cain
Admin
26 days ago

…to test quantum gravity we’d have to be able to measure differences 10 to the 23rd power times smaller than an atom.

A colleague of mine once said that if turns out quantum space-time is how things are, the most disappointed people will be the mathematicians. All that effort, and the real numbers turn out to be just a useful approximation…

russell
russell
25 days ago

AI and LLMs do not experience the existential weight of decisions. At least as far as any of us can tell, and I think our sense of that is correct. They can reason about things that can be empirically measured, and they can present that reasoning in language that they have been taught is congenial to humans.

But that isn’t the same as understanding the subjective human experience of anything they reason about or parrot back to us.

They are useful tools, profoundly in some cases. But all of the above is why I am opposed to giving them decision making power in any but the most mechanical contexts. And humans should always have the power to override.

They’re great at information. That’s enormously useful. And that’s about it, as far as I can tell.

Hartmut
Hartmut
25 days ago

Unfortunately, certain circles like exactly those aspects. An AI is amoral and not a legal person (or in the gray zone). I think under Rumsfeld the Pentagon seriously discussed the potential of autonomous killer drones as a tool to commit war crimes while escaping responsibility*. One cannot court-martial or punish a computer program and the hardware it uses. And if it acted ‘autonomously’, those who sent it out could not be held responsible either since of course no one ordered the specific war crime to be committed but the killer drone did it on its own. At best the programmer could serve as a scapegoat. Iirc when those discussions became public and there were strong hints that courts would not necessarily swallow that line of reasoning, the topic got dropped. The current gang would clearly not care and ICE would love it (except for those officers of course who would miss the personal bodily abuse of power and would be envious of the machines).

*Yeah, the same gang that tried actively to recruit socio/psychopaths for Iraq and only stopped it when someone told them that these guys would feel no inhibition to target their own superiors too.

CharlesWT
CharlesWT
25 days ago

“Now consider warbots. Since self-preservation would not be their foremost drive, they would refrain from firing in uncertain situations. Not burdened with emotions, autonomous weapons would avoid the moral snares of anger and frustration. They could objectively weigh information and avoid confirmation bias when making targeting and firing decisions. They could also evaluate information much faster and from more sources than human soldiers before responding with lethal force. And battlefield robots could impartially monitor and report the ethical behavior of all parties on the battlefield.”

Let Slip the Robots of War: Lethal autonomous weapon systems might be more moral than human soldiers.

nous
nous
25 days ago

Charles, that Reason op ed is more than a decade old, and imagines a sort of warfighting that is rapidly becoming a historical artifact. It also imagines command structures and decision trees that are at odds with how drone warfare has been implemented. We are not heading towards a future full of infantry bots. What we are heading towards is a future of targeting infrastructure, economies, and logistics, and the AI guided autonomous bots will basically act as kamikaze troops and kamikaze interceptors.

I’m also dubious about the rational choice argument where removing the self-preservation motive leads to clearer, more objective decisions. What we have actually seen – especially with the current clown show in the US, is a preference for low-risk/high potential reward tactics based on very limited intel. Low-cost systems with no families create no domestic backlash, and autonomous decision making provides a lot of room for moral deniability – which is already a huge problem.

Do you actually find that Reason argument compelling, or is it just something that appeals narratively to a sense of science fiction imagination?

nous
nous
25 days ago

Had my last appointment and x-ray with orthopedics. Got the green light to chuck the brace and resume normal activities. Still working to get back strength and range of motion, but the break is no longer an issue.

Tom H
Tom H
12 days ago

@Pro Bono:

We established a bridge between these two complementary spheres by fine-tuning a Gemini model to automatically translate natural language problem statements into formal statements, creating a large library of formal problems of varying difficulty.

(https://deepmind.google/blog/ai-solves-imo-problems-at-silver-medal-level/)

That’s AlphaProof including some dependence on LLMs two years ago, and Zahavy could be talking about subsequent work. And, in fact, the next article on Zahavy’s blog (https://www.tomzahavy.com/projects/alphaproof) says “Starting from a pre-trained LLM exhibiting proficiency in mathematics, AlphaProof embarked on a lifelong Reinforcement Learning journey”.

I’m not sure how you’re finding the argument that AlphaProof isn’t an LLM? Lots of modern “AI” is (LLM + tooling and harness).

Pro Bono
Pro Bono
12 days ago

To quote from your source: It couples a pre-trained language model with the AlphaZero reinforcement learning algorithm, which previously taught itself how to master the games of chess, shogi and Go.

The reinforcement learning algorithm is not a LLM.

Last edited 12 days ago by P B