Apple’s new study shows that advanced AI reasoning models like OpenAI’s o3, Anthropic’s Claude, and DeepSeek’s R1 fail completely when problems become too complex.
These models use “chain-of-thought” reasoning to solve puzzles step-by-step, but beyond a certain difficulty level, their accuracy collapses sharply, and they stop reasoning effectively, even with enough computing resources or given solutions. This suggests that current AI reasoning is mostly pattern matching, not true understanding.
The findings challenge claims that today’s AI models are close to human-like general intelligence and highlight fundamental limits in current approaches.