← Latest briefing

Technology

AI models show marked progress on puzzles but struggle with spatial reasoning and scaling complexity

Multiple benchmark studies demonstrate that despite rapid gains on specific games, large language models continue to falter on visual tasks, subtle variations, and multi-step logic.

The short version

  • Artificial intelligence models have shown rapid performance gains on various puzzles and games, but research shows they still face clear cognitive limitations compared to humans.
  • Models frequently fail at 3D spatial reasoning, visual abstract puzzles, and reasoning tasks where minor prompt adjustments alter memorized training data.
  • Performance for top models deteriorates when logical complexity increases, such as managing six or more elements in classical river-crossing or disk-stacking problems.
  • It remains debated whether failures on complex logic tasks point to fundamental flaws in language model reasoning or expected error scaling under heavy cognitive loads.

Key facts

  • Puzzles and games have been used to evaluate artificial intelligence performance since the 1950s.[MIT Technology Review]
  • A late 2024 study by Columbia University researchers found top AI models initially solved only 18 percent of New York Times Connections puzzles, though some models achieved near-perfect scores by early 2025.[MIT Technology Review]
  • Research from Apple revealed that large language models easily handle simple versions of Tower of Hanoi and river-crossing problems but begin to fail when the task involves six or more disks or individuals.[MIT Technology Review]
  • Studies by researchers at Google, the University of Illinois Urbana-Champaign, the University of Washington, Stanford University, and the Allen Institute for AI demonstrated that models struggle when presented with modified versions of memorized puzzles or increasingly complex logic grids.[MIT Technology Review]

What remains uncertain

  • Commentators disputed whether performance drops on highly complex logic puzzles indicate a unique structural weakness in LLM reasoning or merely a standard accumulation of errors as task difficulty increases.[MIT Technology Review]

Sources