SciTech Pulse
AI

AI Models Ace Some Puzzles but Still Fail Basic Spatial Reasoning Tests

AI systems have improved dramatically on some classic puzzles in just months, but researchers find they still struggle with mental rotation, subtle riddle variations and abstract visual reasoning that most humans find…

Puzzles and games have tested artificial intelligence since the field's earliest days: the term 'machine learning' itself was popularised in a 1959 article by IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the board game Go later became well-known AI testbeds. Judged on puzzle-solving alone, AI is improving fast: in late 2024, researchers at Columbia University found that even the best models could solve only 18% of the New York Times' Connections puzzles, but by early 2025 some models were solving them near perfectly.

Spatial reasoning remains a weak spot. In tests, which ask whether different images show the same object from different angles, today's AI models still fail badly even though they can analyse visual inputs. They remain unable to manipulate 3D objects the way human spatial thinkers such as architects can.

Large language models' vast memories can also work against them. A 2024 study by researchers from Google and the University of Illinois Urbana-Champaign tested models on slight variations of 'Knights and Knaves' logic puzzles, in which some characters always tell the truth and others always lie. Models trained on one version often missed small differences in a new one, the study found, and answered from memory instead of reasoning it out.

Abstract visual puzzles that ask for a general rule from a handful of examples, in a benchmark called , are another weak point. Models perform better when a puzzle grid is given as a string of numbers rather than as an image. Research has found that even when they answer correctly, they often rely on convoluted, non-generalisable rules rather than the simple visual concepts humans use.

Terms explained

The story so far

  1. AI Observatory Finds Company Usage Reports Miss Much of Real-World AI Use
  2. AI Agents Still Can't Do Open-Ended Research, Princeton Study Finds
  3. AI Language Models Mirror Users' Political Views, Brazilian Study Finds
  4. NEC Begins Building an AI Foundation Model for Underwater Acoustics
  5. Over 30% Changed Their Stance on Moral Dilemmas After an AI Counterargument, Kobe Study Finds
  6. AI Mines 448 Research Papers to Discover New Heat-Stable, Lead-Free Materials
  7. AI Models Ace Some Puzzles but Still Fail Basic Spatial Reasoning Tests
#artificial intelligence#puzzles#ARC-AGI#spatial reasoning#benchmarks
Rate this story

Related stories