MIT Technology Review is reporting that puzzles and games have been fundamental to artificial intelligence development since its inception, serving as crucial test beds for model advancement. It said that while AI's puzzling skills are improving rapidly, with some models solving New York Times Connections puzzles near perfectly by early 2025, significant weaknesses persist.
The outlet noted that despite advances, current AI models often fumble with subtle changes in classic riddles and struggle particularly with visual puzzles. Spatial reasoning remains a major domain where humans hold a considerable advantage; language models, even with visual input capabilities, perform poorly on mental rotation problems and cannot manipulate 3D objects like human spatial thinkers.
MIT Technology Review said that while frontier large language models possess extraordinary memories from extensive training, this can be a liability. When puzzles closely resemble training data, models may overlook key differences, as seen in a 2024 study involving variations of Knights and Knaves puzzles and tests like SimpleBench. AI also struggles with abstract and visual reasoning in two dimensions, such as with ARC-AGI problems, often using non-generalizable rules even when correct.
The publication highlighted that humans also have cognitive foibles not shared by AI, with psychologists designing problems where humans give knee-jerk answers while models respond deliberately. It also reported that LLMs can ace simple versions of puzzles like the Tower of Hanoi and river-crossing problems, but begin to falter as complexity increases, typically when the number of elements hits six or higher. Similarly, LLMs struggle with logic grid puzzles as complexity piles up.
Full Article: AI models flub these intelligence tests. Can you fare any better?