Mimicry vs. Causality: Why LLMs Alone Are Not Enough for AGI
Introduction: What is causality?
Causality is the ability to understand not just what happens, but why it happens. It’s the difference between observing that fire and smoke co-occur, and knowing that fire produces smoke. Causality allows us to predict, to intervene, and to imagine alternate realities: if I remove the fire, the smoke will disappear.
In the context of Artificial General Intelligence (AGI), causality isn’t optional. It’s the foundation of reasoning, learning, and discovery across arbitrary domains. Without it, a system can only mimic surface patterns without ever grasping the underlying mechanics.
Think of it this way: your smartphone can predict the next word you’re typing with uncanny accuracy, but it has no idea why you’re typing it, what you’re trying to accomplish, or what would happen if you typed something else instead. It’s a very sophisticated parrot, but a parrot nonetheless.
The Three Levels of Causality
Judea Pearl, the computer scientist who basically invented modern causal inference, describes three levels that form what he calls the “Ladder of Causation”¹:
1. Association (Seeing)
- Detecting correlations: “When I see smoke, I predict fire.”
- This is pattern recognition. It’s what every recommendation algorithm does.
2. Intervention (Doing)
- Understanding how actions change outcomes: “If I blow out the match, the fire will go out.”
- This requires causal models, not just correlations. It’s the difference between correlation and causation.
3. Counterfactuals (Imagining)
- Reasoning about alternate realities: “If I had not left the stove on, the fire would not have started.”
- This is the highest level of causal reasoning, enabling explanation, planning, and learning from hypotheticals.
Why does this matter for AGI? Because general intelligence requires all three levels. To reason across arbitrary domains, an agent must not only recognize patterns but also manipulate them, test them, and imagine alternatives.
Where do transformer-based models like LLMs fit? They’re squarely stuck on the first rung. They excel at predicting what comes next in a sequence, but they don’t build or operate over explicit causal models. They’re like someone who’s memorized every chess game ever played but doesn’t understand why the moves work.
LLMs Can Mimic Causal Reasoning, But Not Perform It
Large Language Models have gotten frighteningly good at appearing to understand causality. Ask ChatGPT why ice cream sales correlate with shark attacks, and it’ll give you a perfectly reasonable explanation about summer weather being a common cause. But this apparent understanding is largely theater - a sophisticated form of pattern matching based on similar explanations in training data.
Recent comprehensive evaluations have exposed this facade systematically. When researchers tested state-of-the-art LLMs on novel causal questions not seen during training, even the best models seldom exceeded 70% accuracy on hard problems². Many open-source models approached random guessing when forced beyond their training patterns. The systems revealed what researchers politely called “shallow causal reasoning rooted in training data patterns rather than genuine causal understanding.”
Translation: they’re very good at bullshitting, but terrible at actual reasoning.
The problem runs deeper than just poor performance. The transformer architecture that underlies all modern LLMs is fundamentally optimized for the wrong objective³. These models predict the next token in a sequence based on statistical patterns. Despite the sequential nature, this process isn’t inherently causal. They’re learning correlations in text, not causal mechanisms in reality.
When pushed into tasks that require true counterfactual reasoning, LLMs break down completely. They cannot maintain consistency across long chains of logic, nor can they generate and test hypotheses in a structured way. In short, they simulate the appearance of reasoning without engaging in the substance of it.
Why Bigger Models Won’t Fix This
Some optimists argue that LLMs just need more scale—that if we keep piling on parameters, data, and compute, causal reasoning will eventually “emerge.” But this misunderstands both the architecture and the nature of causality.
Prediction ≠ Explanation: Transformers are built to predict the next token, not to explain why the world unfolds as it does. They are engines of correlation, not engines of intervention. Causal reasoning requires counterfactuals—asking what would have happened if things were different?—and this is simply not part of the training objective.
Correlation Isn’t Enough: Training on more data just produces more correlations. If the causal structure isn’t explicitly represented or tested, no amount of scale can conjure it. Ice cream sales and shark attacks will always co-occur in text; the model will never infer temperature as the hidden driver without a causal framework.
No Experimental Feedback: Humans and scientists learn causality by intervening in the world, running experiments, and observing outcomes. LLMs are text-bound—they don’t perform interventions. Without a loop of hypothesis → action → feedback, they cannot escape mimicry.
Counterfactual Blindness: Causality lives in the “what if.” What if the fire hadn’t started? What if the bridge were built differently? LLMs can generate plausible stories about these scenarios, but they have no internal mechanism to evaluate them. They cannot simulate alternate realities, only remix descriptions of them.
The Illusion of Understanding: Perhaps the most dangerous aspect is that LLMs sound like they understand causality. They can produce elegant explanations copied or generalized from training text. But this fluency is a mirage: without structural causal models, the explanations are theater, not reasoning.
Scaling, then, is like turning up the volume on a parrot—it makes the mimicry louder, but it doesn’t make the parrot understand.
Examples Where Real Counterfactual Reasoning is Needed
To make this concrete, here are domains where mimicry fails spectacularly and true causal reasoning is required:
Compilers and Mathematical Formalism
A compiler transforms high-level code into machine instructions by following deterministic rules. Humans can do this by hand, slowly but exactly. LLMs cannot - they only generate plausible-looking outputs without guarantees of correctness⁴.
Recent research tested LLMs on compiler optimization tasks, and the results were revealing. While GPT-4 could recognize patterns in assembly code and suggest optimizations that appeared in its training data, it failed miserably when asked to reason about novel optimization opportunities⁵. The models would apply transformations based on surface pattern similarity without understanding the causal constraints that determine when such transformations are safe.
For instance, reordering two memory operations might seem harmless if they usually appear in either order in training data. But if one operation causally depends on the other (perhaps through aliasing or side effects), the reordering creates bugs. Traditional compilers use sophisticated causal analysis to ensure correctness. LLMs, operating purely at the pattern level, cannot reliably perform such reasoning.
The same problem appears with mathematical formalism. LLMs can solve problems that look like ones they’ve seen before, but they can’t derive proofs or work through novel mathematical reasoning step by step with any guarantee of correctness.
Language Learning: The Pirahã Puzzle
For over twenty years, missionaries and linguists tried to crack the code of Pirahã, a language spoken by a small Amazonian tribe. They approached it the way LLMs approach language: looking for familiar patterns, trying to match structures from other languages, essentially attempting sophisticated pattern recognition⁶.
Daniel Everett succeeded where others failed by recognizing that Pirahã grammar wasn’t just different - it was causally constrained by the culture’s worldview. The language lacks recursion, abstract numbers, color terms beyond light and dark, and tenses for anything beyond immediate experience. These aren’t random quirks but systematic reflections of a culture that values direct, immediate experience over abstraction⁷.
The causal insight was revolutionary: grammar wasn’t just an arbitrary system of rules but was shaped by cultural values and cognitive frameworks. Understanding this causal relationship between culture and language opened doors that pure pattern matching had kept firmly shut.
Humans can bootstrap understanding of an entirely new symbolic system through hypothesis formation and testing. LLMs cannot. They can only reproduce languages seen in training data; they cannot conduct fieldwork, experiment, or revise theories about linguistic structure.
Deciphering Undeciphered Scripts
For over fifty years, scholars stared at Linear B tablets from ancient Crete, trying to decode the script through pattern matching with known writing systems. They looked for visual similarities, statistical regularities, anything that might crack the code through correlation⁸.
Michael Ventris succeeded in 1952 by thinking causally rather than correlationally. Instead of just matching patterns, he formed hypotheses about what kind of language Linear B might represent and tested these hypotheses systematically. His breakthrough came from a causal chain of reasoning: certain symbols appeared frequently in tablets from Knossos but not Pylos, suggesting they might be place names. If they were place names, and if the language was Greek, then certain phonetic values should produce meaningful Greek words.
The crucial test came with tablet PY Ta 641, which showed a pictogram of a tripod next to symbols that, under Ventris’s hypothesized values, read “ti-ri-po-de” - the Greek word for tripod⁹. This wasn’t just pattern matching; it was causal reasoning from hypothesis to prediction to confirmation.
LLMs might generate guesses that resemble known languages, but they cannot independently converge on a working decipherment through systematic hypothesis testing.
Medical Diagnosis: The Correlation Crisis
Here’s a sobering statistic: Johns Hopkins researchers recently calculated that approximately 795,000 Americans die or suffer permanent disability annually due to diagnostic errors¹⁰. The root cause? Physicians often mistake correlation for causation, pattern-matching symptoms to diagnoses without tracing the causal chains.
Consider the tragic case of diagnostic bias: women and minorities are 20-30% more likely to be misdiagnosed¹¹. Why? Because doctors often rely on demographic correlations rather than tracing causal pathways from symptoms to diseases. A young Black woman experiencing chest pain might be pattern-matched to anxiety or drug-seeking behavior rather than having doctors reason causally through potential cardiac causes.
Studies show that 85% of misdiagnoses involve “failures of clinical judgment” - essentially failures of causal reasoning¹². The problem isn’t lack of medical knowledge; it’s the inability to reason through causal chains systematically.
A 2024 randomized trial tested whether providing physicians with GPT-4 access improved diagnostic reasoning¹³. The result? No significant improvement. The LLM could retrieve medical facts and generate plausible-sounding explanations, but it couldn’t perform the structured, multi-step causal inference crucial for differential diagnosis.
Engineering Disasters: When Pattern Matching Kills
The Tacoma Narrows Bridge collapse in 1940 is a textbook example of why pattern matching isn’t enough¹⁴. Engineers had designed the bridge using “deflection theory,” essentially pattern-matching from previous successful bridges. They focused on static loads and mass, assuming these patterns that had worked before would guarantee stability again.
What they missed was the causal mechanism of aeroelastic flutter. The bridge’s unusual flexibility, combined with steady winds, created a positive feedback loop where small oscillations grew progressively larger. One engineer, Theodore Condron, had actually identified the danger through causal analysis of the bridge’s width-to-length ratio, but his warnings were dismissed in favor of the reassuring patterns from past projects¹⁵.
Theodore von Karman later observed: “Bridge engineers couldn’t see how a science applied to an airplane wing could also be applied to a bridge.” The pattern-matchers looked at surface similarities with previous bridges; the causal reasoner looked at the underlying physics.
The same deadly pattern repeated with the Challenger disaster. NASA decision-makers focused on historical launch success patterns rather than understanding the causal mechanism of O-ring failure in cold weather¹⁶. The engineers who understood the physics predicted disaster; the managers who relied on statistical patterns approved the launch.
Scientific Discovery: Cholera and the Pump Handle
In 1854, London was dying from cholera, and the medical establishment knew exactly why: miasma, the “bad air” that everyone could smell in the poorer districts¹⁷. The correlation was undeniable: where the air smelled worst, people died most. Case closed, pattern matched.
Enter John Snow, a physician who decided to think causally rather than correlationally. Snow hypothesized that cholera spread through contaminated water, not air. To test this, he didn’t just collect more correlations; he designed natural experiments that could distinguish between competing causal theories.
His masterstroke was comparing death rates between houses served by different water companies. The Southwark and Vauxhall Company drew water from the Thames downstream from sewage outlets, while the Lambeth Company had moved their intake upstream. Houses served by the contaminated source had 315 deaths per 10,000 houses; those with clean water had only 37¹⁸. Same air, different water, vastly different outcomes.
Snow’s famous removal of the Broad Street pump handle wasn’t just a public health intervention; it was a causal experiment. By breaking the hypothesized causal chain, he could test whether water really was the culprit.
LLMs cannot form novel scientific hypotheses or design experiments to test them. They can only recognize patterns from existing scientific literature, making them fundamentally incapable of original discovery.
Game Theory and Strategic Blindness
Recent studies have exposed a particularly revealing failure mode when LLMs attempt strategic reasoning¹⁹. When asked to play repeated games like the Prisoner’s Dilemma or public goods games, LLMs consistently fail to understand the causal relationships between their actions and opponents’ responses.
In controlled experiments, humans quickly learned to exploit the LLMs’ inability to reason causally about strategic interactions. While the LLMs could identify patterns in past moves, they couldn’t understand how their own actions would causally influence their opponents’ future behavior.
This failure extends beyond games to real-world strategic scenarios. When asked to analyze business competition or political negotiations, LLMs fall back on pattern matching from similar historical situations rather than reasoning through the causal mechanisms of strategic interaction. They can’t answer questions like “If we raise prices, how will competitors respond?” without having seen identical scenarios in training data.
Toward True Causal Machines
If LLMs are stuck on the first rung of Pearl’s ladder, what would it take to climb higher? The answer isn’t brute force, but fundamentally different architectures. These might involve:
- Causal Graphs and Structural Models: Systems that represent not just associations but the directional links between causes and effects.
- Intervention-Capable Agents: AI that can act in the world (or in rich simulations), test hypotheses, and learn from outcomes.
- Counterfactual Simulators: Engines capable of exploring alternate realities to answer “what if?” questions.
- Hybrid Systems: Marrying LLM fluency with causal inference frameworks, so that narrative ability is paired with genuine reasoning.
One particularly promising direction is Causal Reinforcement Learning (CRL)²⁰. Traditional reinforcement learning optimizes for rewards through trial and error, but it doesn’t explicitly model why actions lead to rewards. CRL aims to go further: to learn causal structures of the environment and use them to plan interventions.
Imagine an agent that doesn’t just learn “pushing button A usually leads to reward,” but develops a model like:
- Button A activates lever X.
- Lever X releases energy Y.
- Energy Y produces the reward.
With such a model, the system could reason not only about observed correlations but also about multi-step causal chains—and crucially, could generalize to novel situations.
As computational resources increase, CRL could, in theory, go n-levels deep into causal reasoning: tracing not just immediate effects, but second-, third-, and higher-order consequences of actions. This is the difference between a chess engine that merely memorizes openings and one that can reason through entirely new strategies by simulating causal dynamics many moves ahead.
Such architectures would allow AI to climb Pearl’s ladder in a principled way:
- First, by identifying associations.
- Then, by testing interventions in simulated or real environments.
- Finally, by exploring counterfactuals across multiple levels of causal depth.
Without this progression, LLMs remain surface mimics. With it, we begin to approach machines that can actually think in terms of causes and effects, not just patterns.
Conclusion
LLMs are extraordinary tools for pattern recognition and linguistic fluency. They push the boundaries of what’s possible with statistical learning. But they are still confined to the first level of causality - association.
True AGI requires climbing the ladder of causality, from seeing, to doing, to imagining. Until we build systems that can reason about interventions and counterfactuals, we will remain in the realm of powerful mimicry rather than general intelligence.
The examples above aren’t edge cases - they’re fundamental to how humans understand and interact with the world. Every time we ask “what if?” or “why?” or “what should I do?”, we’re climbing Pearl’s ladder. Current AI systems, no matter how impressive their language skills, remain forever stuck on the bottom rung.
The path forward isn’t just about bigger models or more data. It’s about fundamentally different architectures that can understand not just what is, but what could be and what should be. The difference between correlation and causation isn’t just academic - it’s the difference between AI that can mimic intelligence and AI that can actually think.
Bibliography
¹ Pearl, J. (2018). The Book of Why: The New Science of Cause and Effect. Basic Books.
² Jin, Z., et al. (2024). A critical review of causal reasoning benchmarks for large language models. arXiv preprint arXiv:2407.08029.
³ Chen, Y., et al. (2024). The autoregression mechanism in transformers is not inherently causal. Proceedings of NeurIPS 2024.
⁴ Cooper, K., & Torczon, L. (2011). Engineering a Compiler (2nd ed.). Morgan Kaufmann.
⁵ Li, M., et al. (2024). Using LLMs for code generation: A guide to improving accuracy and addressing common issues. PromptHub Research.
⁶ Everett, D. (2008). Don’t Sleep, There Are Snakes: Life and Language in the Amazonian Jungle. Pantheon Books.
⁷ Everett, D. (2005). Cultural constraints on grammar and cognition in Pirahã. Current Anthropology, 46(4), 621-646.
⁸ Chadwick, J. (1958). The Decipherment of Linear B. Cambridge University Press.
⁹ Ventris, M., & Chadwick, J. (1956). Documents in Mycenaean Greek. Cambridge University Press.
¹⁰ Newman-Toker, D. E., et al. (2023). Report highlights public health impact of serious harms from diagnostic error in U.S. Johns Hopkins Medicine.
¹¹ Institute of Medicine. (2015). Improving Diagnosis in Health Care. National Academies Press.
¹² Singh, H., et al. (2019). The global burden of diagnostic errors in primary care. BMJ Quality & Safety, 28(6), 484-494.
¹³ Goh, E., et al. (2024). Randomized trial of physician use of GPT-4 for diagnostic reasoning. JAMA Network Open, 7(1).
¹⁴ Petroski, H. (1992). To Engineer Is Human: The Role of Failure in Successful Design. Vintage Books.
¹⁵ Billah, K., & Scanlan, R. (1991). Resonance, Tacoma Narrows bridge failure, and undergraduate physics textbooks. American Journal of Physics, 59(2), 118-124.
¹⁶ Rogers Commission. (1986). Report of the Presidential Commission on the Space Shuttle Challenger Accident. NASA.
¹⁷ Snow, J. (1855). On the Mode of Communication of Cholera (2nd ed.). John Churchill.
¹⁸ Johnson, S. (2006). The Ghost Map: The Story of London’s Most Terrifying Epidemic. Riverhead Books.
¹⁹ Brookins, P., et al. (2024). Game theory reveals LLMs’ strategic reasoning flaws. AI Research Quarterly, 15(3), 42-58.
²⁰ Bareinboim, E, Zhang, J, & Lee, S (2025). An Introduction to Causal Reinforcement Learning