AI models flub these intelligence tests. Can you fare any better?

Summarized from technologyreview.com


The article discusses the role of puzzles and games in evaluating the progress of artificial intelligence (AI) models, tracing this practice back to early AI development, such as Arthur Samuel’s checkers-playing algorithm in 1959. It highlights rapid improvements in AI’s puzzle-solving abilities, citing a Columbia University study where AI models’ success rates on New York Times Connections puzzles surged from 18% in late 2024 to near-perfect performance by early 2025. However, the article also notes persistent weaknesses in AI, particularly with subtle variations of classic riddles and visual puzzles, suggesting these areas offer insight into the divergent strengths and limitations of machine versus human cognition.

The piece provides several examples illustrating these differences. It points out that while large language models (LLMs) excel at recalling vast amounts of factual information, this can lead to errors when puzzles closely resemble training data, as seen in a 2024 study involving “Knights and Knaves” puzzles. Additionally, the article describes how AI struggles with spatial reasoning tasks, such as mental rotation problems, and abstract rule inference in benchmarks like ARC-AGI, often employing non-generalizable strategies compared to humans’ intuitive approaches. The text also acknowledges human cognitive biases that can be exploited in certain problem sets, where AI may outperform humans by avoiding intuitive but incorrect answers. Finally, it notes scalability issues for AI in complex logical puzzles, such as the Tower of Hanoi and logic grid puzzles, where performance declines as problem complexity increases.