Google DeepMind Researcher Says Current LLMs Cannot Recreate Human Scientific Genius

A new position paper from Google DeepMind researcher Tom Zahavy argues that current large language models cannot independently make the creative conceptual leap behind major scientific breakthroughs. Titled LLMs Can’t Jump, the paper examines whether modern artificial intelligence could have invented General Relativity using only the scientific knowledge available to Albert Einstein. Its conclusion is that current systems could potentially complete much of the mathematical work, but would struggle to invent the original principles needed to begin.

Zahavy separates reasoning into induction, deduction, and abduction. Induction identifies patterns and derives general rules from accumulated examples, while deduction applies established rules to reach logically verifiable conclusions. Abduction is different because it proposes a new explanation, concept, or cause when the available evidence is limited. The paper argues that large language models have become highly capable at induction and are making rapid progress in deduction, as demonstrated by systems such as AlphaProof, but they remain dependent on concepts, objectives, and symbolic frameworks that already exist.

Einstein’s development of the equivalence principle is presented as the central example. Instead of discovering General Relativity by processing a large collection of observations, Einstein imagined the physical experience of a person falling freely or standing inside an accelerating elevator. This thought experiment allowed him to connect gravity and acceleration before sufficient experimental evidence existed to confirm the theory. According to Zahavy, an artificial intelligence system given Einstein’s principles could plausibly derive their mathematical consequences, but a conventional language model would not independently generate those principles from physical intuition.

"The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought."
— Quote by: Albert Einstein

The paper does not claim that artificial intelligence will permanently remain incapable of scientific invention. Instead, it argues that future architectures must move beyond text and passive visual prediction. Zahavy proposes physically consistent multimodal world models that allow artificial intelligence agents to interact with simulations, manipulate environments, test counterfactual scenarios, and connect abstract symbols with physical consequences. That direction closely aligns with DeepMind’s work on Genie, which creates controllable interactive environments rather than conventional generated videos. As previously examined through Project Genie and playable world research, these systems remain experimental but could provide training environments for more grounded artificial intelligence agents.

The distinction is also important because this is a theoretical position paper, not experimental proof that all current or future artificial intelligence systems face an absolute limit. DeepMind already uses systems such as AlphaEvolve to discover and optimize algorithms within measurable frameworks, but Zahavy argues that optimization inside an established system is fundamentally different from inventing an entirely new scientific framework without a clear objective or error signal.

The paper offers a more practical interpretation of the human versus artificial intelligence debate. Current models can accelerate research, search enormous knowledge bases, generate code, verify proofs, and explore thousands of possible solutions. What remains uncertain is whether they can independently decide which unexplored question matters, invent the conceptual framework needed to investigate it, and recognize a breakthrough before the supporting evidence exists.

The same limitation applies to gaming and creative technology. Artificial intelligence may generate environments, assets, dialogue, code, and design variations, but speed does not automatically produce creative intent, emotional meaning, or a coherent vision. Human creators remain responsible for defining why an experience should exist and what it should communicate.


Will physically grounded world models eventually give artificial intelligence genuine creative intuition, or will major scientific breakthroughs continue to depend on human experience?

Share
Angel Morales

Founder and lead writer at Duck-IT Tech News, and dedicated to delivering the latest news, reviews, and insights in the world of technology, gaming, and AI. With experience in the tech and business sectors, combining a deep passion for technology with a talent for clear and engaging writing

Previous
Previous

Double Fine Cuts 23 Jobs Weeks After Regaining Independence From Xbox

Next
Next

WUCHANG Fallen Feathers Sequel Reunites 505 Games With Its Original Director