Sitemap

Beyond the Puzzle Box: Why The “Illusion of Thinking” Paper Misreads the Applied AI Revolution

5 min readJun 14, 2025

--

Press enter or click to view image in full size

A recent paper, “The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity,” by Shojaee et al., has made a provocative claim: that the apparent “thinking” of advanced AI is largely a mirage. The authors present a methodical study using controlled puzzles, and for this, they should be commended. Their work provides valuable data on how Large Reasoning Models (LRMs) perform on specific, algorithm-heavy tasks.

While the data is interesting for the limited scenario of e.g., tower of hanoi problem, the paper’s grand conclusion seems like an overreach, built on a narrow methodology, a misinterpretation of model behavior, and a failure to acknowledge the true scope of modern AI reasoning.

Here we discuss the paper’s core claims and inspect them to see if they hold up with more scrutiny.

The Problem with Puzzles: A Flawed Yardstick for Intelligence

The paper’s entire thesis rests on the performance of models on a handful of puzzles: Tower of Hanoi, Checker Jumping, River Crossing, and Blocks World. In Figure 4, the authors show a clear trend: as the complexity of these puzzles increases (e.g., more disks in Tower of Hanoi), model accuracy plummets to zero. They interpret this “complete accuracy collapse” as a fundamental failure of reasoning.

This conclusion mistakes a specific architectural limitation for a universal reasoning failure.

These puzzles primarily test algorithmic planning and perfect procedural fidelity. They have defined rules and optimal paths, making them excellent tests for classical AI search algorithms. However, they are a questionable proxy for the ambiguous, gray-area spectrum, of multifaceted nature of human and advanced AI reasoning.

By focusing on a domain where current sub-symbolic LLMs are known to be weaker — perfect, step-by-step symbolic execution over long sequences — the paper essentially sets them up to fail.

The collapse in accuracy on a 15-disk Tower of Hanoi problem doesn’t prove the model can’t “think.” Thinking or reasoning models may not think or reason like we do as humans, but in most cases they are getting the job done, like successfully completing Olympiad level math problem.

It proves that its neural, pattern-matching architecture is not optimized to be a flawless stack-based state machine. It’s like judging a brilliant novelist’s intelligence by their inability to manually calculate the 1,000th digit of pi. The test doesn’t fit the subject or the variety of application that are enabled through thinking models.

Debunking the Inefficient Mind: A Misreading of AI Behavior

The paper delves deeper, analyzing the “reasoning traces” or “thinking tokens” generated by the models. It highlights two phenomena as evidence of flawed thinking:

  1. The “Overthinking” Phenomenon: In Figure 7a, the paper shows that even for simple problems, models often find the correct solution early in their thought process but continue to explore incorrect alternatives. This is framed as inefficient “overthinking.”
  2. The “Counterintuitive” Decline in Effort: Even more central to their argument is the finding in Figure 6, where at very high complexities, the models’ “reasoning effort” (the number of thinking tokens) decreases instead of increases. The paper calls this a “fundamental scaling limitation.”

The Ultimate Test: The Failure of the Provided Algorithm

Perhaps the most damning piece of evidence in the paper — and the most revealing to refute — is presented in Figure 8. Here, the authors gave the models the explicit, step-by-step recursive algorithm for solving the Tower of Hanoi. Shockingly, performance did not improve. They conclude this demonstrates a core limitation in “logical step execution.”

This doesn’t prove the models can’t reason; it proves they are not traditional code interpreters. For an LLM, an algorithm in a prompt is just more text. It doesn’t automatically compile and execute it. The model has to learn to map this declarative knowledge (the text of the algorithm) to procedural action, a notoriously difficult challenge in AI known as neuro-symbolic integration. This failure highlights an architectural characteristic — that

LLMs are not symbolic logic engines —

rather than an absence of a capacity for thought.

The Real Test: Reasoning Beyond the Puzzle Box

If the paper’s puzzles are a poor measure, what’s a better one? Consider problems that require not just following rules, but inventing them: Math Olympiad problems.

The ability of models like Google’s Gemini Pro 2.5 to solve these problems is a powerful counter-argument to the “illusion” thesis. Unlike the Tower of Hanoi:

  • Olympiad problems are non-routine. There is no pre-defined algorithm. The model must devise a novel strategy.
  • They require deep conceptual insight. The model must understand abstract concepts in number theory, geometry, and algebra, not just manipulate tokens.
  • They demand creative problem-solving. Solutions often involve an “aha!” moment — a clever substitution, an elegant construction, or a unique framing of the problem.
  • They require rigorous proof. The model must generate a logically sound, step-by-step argument, demonstrating an incredible internal consistency.

If a model can devise a creative proof for a problem from the International Mathematical Olympiad, the claim that its thinking is an “illusion” because it fails to flawlessly execute the 2¹⁵ — 1 steps of a Tower of Hanoi puzzle becomes untenable. It demonstrates that the paper was measuring with a yardstick when it needed a barometer.

Conclusion: The Illusion is in the Lens, Not the Light

The “Illusion of Thinking” paper provides a valuable, if limited, snapshot of how certain AI models handle procedural, algorithmic tasks and speaks well to the quite limited use-case domain (e.g., tower of hanoi) in which they are not super effective. Its data is useful for understanding the current limitations of LLMs as symbolic execution engines.

But its conclusion is a profound mischaracterization. The real story of AI reasoning is not one of collapse in the face of simple puzzles, but of astonishing emergence in the face of complex, abstract, and creative challenges. The thinking is not an illusion; the illusion is believing that a simple puzzle box could ever contain the full measure of this rapidly expanding intelligence.

--

--

Ali Arsanjani
Ali Arsanjani

Written by Ali Arsanjani

Director Google, AI | EX: WW Tech Leader, Chief Principal AI/ML Solution Architect, AWS | IBM Distinguished Engineer and CTO Analytics & ML